No File, No Score, No Loan, No Way to Start One
Tens of millions of American adults have no credit file or too thin a file to generate a score. The system that decides whether they can borrow requires a history of borrowing they were never able to start.
Three Different Situations
People group three different conditions under one phrase: bad credit. They are not the same problem and mixing them up is where most of the public conversation goes wrong
invisible credit It means exactly what it sounds like. There is no file in any of the three main offices. No one has ever reported anything about this person's debt because this person has never borrowed or has only done so in ways that no one reports
Not scoring It's one step up and one step less than usable. One file exists. It just doesn't contain enough to generate a score based on conventional models. An account opened last month can do it. So can several accounts that went silent ten years ago
Subprime It is different in type not just degree. There is a score and it is low because it reflects the real story: late payments high utilization a pattern that a model can actually read
Here's the distinction that matters. Subprime mortgages are a judgment about behavior. The other two are not judgments at all. They are an absence of information. When a lender rejects credit from an invisible applicant it is not saying that the applicant is risky. That is to say he has no idea and no idea is exactly treated as high risk by a system built to value risk not uncertainty
| Status | What the lender sees | Typical result |
|---|---|---|
| invisible credit | No file | Automatic rejection on most systems. |
| Not scoring | Insufficient or obsolete data | Manual review or rejection |
| Subprime | A low score | Passed at a high rate |
Who It Affects
Consumer financial regulators have put ballpark numbers on this and the raw coverage is actually working. Research on that side of the government has estimated that about one in ten American adults has invisible credit plus another several percent who can't be scored. If we add the two together we're talking about tens of millions of people though I wouldn't defend the third decimal of any of these estimates in a room. Population surveys of this type rely on matching bureau records with census data and thecomparison itself is imperfect
What is not disputed is that the distribution is uneven. Young adults show up at much higher rates for the obvious reason that they have not yet had time to build up a file. The same goes for recent immigrants who may have excellent credit in another country to which it simply does not transfer. So do residents of low-income neighborhoods and specifically black and Hispanic consumers. That grouping is what turns a data engineering problem into a policy problem because a missing file rarely arrives.alone.It appears accumulated on top of any other restrictions that the same person already has
A thin file is not a fact about a person it is a fact about which of their financial behaviors someone decided to report. Paying rent for fifteen years produces no record. Missing two credit card payments produces a permanent one
The Reporting Asymmetry
The credit system was created to record a limited portion of financial life: loans credit cards and similar extensions of credit reported by providers who volunteer to participate. No one is required to report anything. The whole apparatus works when companies choose to send data to a bureau because it is useful to them
That design choice quietly excludes the largest most regular payment most households make. Rent is the largest monthly obligation for tens of millions of renters and it's almost never reported. Utilities phone bills and insurance premiums follow the same pattern. Pay them on time for years and nothing will happen to your record. So many are lost that the account goes to collections and suddenly it's reported immediately and permanently
Sit with that for a second because asymmetry is the whole story in miniature. A decade of on-time rentals produces zero data points. A bad run with a cable company produces a collections entry that can drag a score down for years. The system was built to detect failures in a few specific categories. It was never built to achieve success in the categories where most people actually succeed
Why This Is a Coverage Problem, Not a Character Judgment
Step away from the individual cases for a second and ask what a credit score actually is. It's not a verdict on someone's character. It's the result of a statistical model trained to predict the probability that a given set of inputs will default over some period usually two years. The model needs inputs. Account age payment history utilization mix of account types all of that. If you remove the inputs the model has nothing to calculate. It's not that the model calculates high risk.for an invisible credit applicant. You literally can't calculate anything at all
That's why I think the system's framework is biased against people without credit and underestimates what's happening even though the outcome is genuinely unfair. A story of pure discrimination implies that the model looked at someone and decided against them. What's actually happening is closer to a coverage gap: the model was never trained and can't evaluate cases that fall outside its input space. That's a limitation of modeling before anything else
Lenders then have to decide what to do with an application that the model cannot qualify and this is where economics not malice comes in. A lender that cannot distinguish an invisible low-risk applicant for credit from a high-risk one faces what economists call adverse selection the same problem that appears in insurance markets when an insurer cannot distinguish healthy applicants from sick ones. If a lender prices a loan to an unqualified population at an average rate applicants who know they are actually fromlow risk and who have other options reject the offer. Applicants who know they are at higher risk and have no better option accept it. The group that actually takes the loan is worse than the group the lender priced it for. Rejecting the entire segment is often the safest and from the lender's narrow perspective the most rational response to that dynamic. It is a bad outcome for the applicant. It is not irrational for the lender
A credit score is not a verdict on character. It is the result of a model that needs inputs to calculate anything and a lender who cannot distinguish a safe borrower from a risky one within a population without a score will often be better off in the strict financial sense by rejecting them all
A Worked Example: Pricing What You Cannot See
Let me put real numbers on the coverage problem because arithmetic works faster than words. Everything here is illustrative round numbers chosen to make the mechanism visible not an actual loan book
Suppose a lender is deciding whether to make 1,000 small personal loans of $2,000 each to a population that cannot qualify for a total of $2,000,000 borrowed. Suppose that if a borrower defaults the lender recovers 40 percent of the loan through collections or guarantees so it loses 60 percent or $1,200 on each loan.default. That 60 percent figure is what insurers call loss given default and I'm setting it to a round number for clarity
Now run two scenarios that differ only in the actual and unobserved default rate for this population
| Scenario | Actual default rate | Defaults per 1,000 loans | Total loss | Loss rate in the book. |
|---|---|---|---|---|
| Scenario A | 5 percent | 50 | 60,000 dollars | 3 percent |
| Scenario B | 15 percent | 150 | $180,000 | 9 percent |
Same 1,000 people same loan amount same recovery assumption. The only thing that changed is a default rate that the lender cannot observe and the required loss reserve triples from 3 percent of the book to 9 percent. A rate built to cover 3 percent of expected losses plus financing costs and a margin will lose money directly in scenario B. A rate built to cover 9 percent would seem expensive and uncompetitive if population were actually the scenario.A
This is exactly the problem a lender with an invisible credit fund faces. There is no file to indicate whether scenario A or scenario B is being analyzed and guessing wrong in either direction is expensive. Suppose you split the difference and price a default rate of 10 percent expecting 1,000 times 0.10 times $1,200 which is equivalent to a $120,000 loss or 6 percent of the book. That price nowis too expensive for anyone who actually belongs in Scenario A so the lowest-risk borrowers in the group those with real alternatives like a collateral or a secured card drop out. Those left with the loan are disproportionately Scenario B borrowers those with fewer options. Realized losses rise toward 9 percent even though the lender set a price of 6 and the loan portfolio performs worse than the assumed price for reasons that have nothing to do with skill.of the lender and yes with who chooses an unattractive price. This is adverse selection with real numbers associated and that is why pricing it a little higher to be safe doesn't really solve the coverage problem. It can make it worse
Case Study: Experian Boost and the Limits of a Bolt-on Fix
The clearest real-world attempt to address this specific gap is Experian Boost which Experian launched in 2019. The mechanism maps directly to reporting asymmetry from some sections up. A consumer connects their bank account and Boost scans it for timely payments on utility bills phone bills and certain streaming subscriptions categories that the traditional office file never captured. It then adds that positive payment history directly to the consumer's Experian file which can elevatea score and in some cases converting a previously ungradeable file into one that can be graded
What I find really useful about Boost as a case study is that it's a clean natural experiment on exactly the coverage problem described above. It doesn't change anyone's underlying financial behavior. The same phone bill paid in the same way on the same day counts or not depending on whether the consumer opted for the product. That's informational asymmetry made visible: a value that already existed made readable to a model that couldn't see it before
It's also a useful lesson in the limits of a bolt-on solution and I want to be honest about what Boost doesn't do. First it requires a bank account to connect so it does nothing for the unbanked who largely overlap with the invisible credit population. It requires the bill to be in the consumer's name so a young adult on a parent's family phone plan an extremely common way for young adults to pay for phone service doesn't generate any data under it.program. And it only helps at the margin for a file that contains almost nothing else. A truly empty file gets one data point not a credit history. Boost is a real improvement on the boundaries of the coverage problem. It's not a half-assed solution
What Has Been Tried
Beyond Boost several other approaches attack the same reporting gap from different angles
Rental reports The services promptly report rent payments to the offices either through a landlord who is already set up to do so or through a third party that independently verifies the payment and submits it. The research on this is really encouraging. It shows real improvements in scoring and more importantly it shows previously non-scorable tenants becoming scorable. The limiting factor is participation not science. Most landlords have no particular incentive to set this up so theCoverage remains patchy
Alternative data Underwriting is the broader category that Boost belongs to: getting information about cash flow directly from a bank account the regularity of income spending patterns the pace of money coming in and going out rather than just relying on what the offices already have. Several lenders have built this into their underwriting and regulators have published guidelines on how to do this without running into fair lending problems which is the topic of the next section
Data authorized by the consumer It's the legal and technical plumbing underneath most of this. It allows an applicant to grant a specific lender temporary access to their own bank account data rather than the data automatically flowing through an office relationship. It's especially useful for someone with a stable income and literally no history of debt since otherwise there is no file to appeal to
There's also a separate category of products secured cards and credit-building loans created specifically to create payment history from scratch when none exists. They work and the mechanics deserve their own treatment elsewhere because this article is about something more specific: why the model can't see these applicants in the first place not what to deliver to them once it can
The Fair Lending Complication
None of this makes alternative data an automatic improvement and the caution that comes with it is real not just a legal throat-clearing
Any variable a lender uses must avoid producing unlawful discrimination and a variable can do so even without discriminatory intent simply by correlating with a protected characteristic close enough to produce what the law calls disparate impactData sources like what college someone attended general purchasing patterns or the make and model of your phone have been questioned on exactly this basis because each can act as a silent surrogate for race national origin or income in ways the lender never intended and may not even realize
Cash flow data tends to have a friendlier reception from regulators and the reason is more instructive than arbitrary. It measures something like the real question the underwriting is trying to answer which is whether this person can repay a loan with their actual income and expenses rather than acting as a proxy for that question like a zip code or purchasing pattern does
Then there is the explainability requirement which is easy to underestimate until you have to build around it. Lenders must give specific individual reasons when they turn someone down. That one rule rules out an entire category of otherwise powerful models those whose predictions can't be traced back to a short list of identifiable factors that a human can read aloud. A model can be highly predictive and still be unusable in this industry if no one can explain in plain language why it said no
Where This Breaks: Alternative Data's Own Blind Spots
I've made the alternative data look like a neat patch on a modeling gap so let me argue against my own framework because the limits are real and specific not just theoretical
The first limit is that every version of this still requires a digital financial footprint to begin with. Cash flow underwriting needs a bank account. Rent reports need a landlord or verification service willing to participate. Boost needs an invoice in your own name and a bank account to connect it. The population most likely to be genuinely and profoundly invisible someone who is unbanked receiving cash payments living in informal or family housing without any lease in their name does not generate any digital trail to which you canGrab any of these products. Alternative data mainly helps people who have thin files. It does much less for people who barely touch the formal financial system
The second limit is selection bias in the evidence itself and this is underestimated in the way these programs are reported. Studies that show that rental reports or Boost-style products increase scores typically study people who chose to participate.People who proactively connect a bank account to a scoring tool or who ask their landlord to start reporting rent are arguably more financially compromised than the average thin-file consumer for starters. Part of the measured improvement may be the treatment. Part of it may simply be who chooses to take the treatment in the first place and disentangling those two is really difficult to get right
The third boundary goes the other way and matters for a different subset of people. Cash flow data works best with stable income. It's much noisier for gig work seasonal work or anyone whose income actually varies month to month which describes many of the same young adults and recent immigrants who are overrepresented among the invisible credits in the first place. More data doesn't always mean a clearer signal. Sometimes it means a more volatile one and a volatile cash flow record can seem riskier for a model.that have almost no record
The fourth limit is the fair lending point from the previous section forward rather than backward. Expanding the set of reportable behaviors also expands the set of variables that could be correlated with a protected characteristic in a way that no one anticipated. Rental payment history is correlated with neighborhood. Neighborhood is correlated in most American cities with race and income for reasons that have nothing to do with credit risk and everything to do with real estate history. Expand the data used for underwritingIt is not a free action. You have to check it for the same disparate impact problem it is trying to solve
As a whole the alternative data reduces the coverage gap. It does not close it and for the most difficult cases those furthest from the formal financial system the figure barely moves
How I Actually Read a Credit Invisibility Story
This is how I use all of this when I read a story about a fintech that claims to serve credit invisibly because that narrative comes up all the time and I think it deserves more skepticism than it usually receives
The first question I ask is which of the three categories at the top of this article the company is actually targeting because marketing rarely says so. Serving the unbanked can mean building a genuinely new coverage path for someone with no track record. It can also mean quietly underwriting subprime and using friendlier language because it works better with investors and regulators. These are extremely different companies using the same phrase and I've learned to ask which one I'm actually looking at before taking the proposition at face value
The second thing I check is what data the product actually requires to work which is really asking who is left out. If a company's alternative data model requires a linked bank account I immediately want to know their answer for someone who is unbanked because that is a real large subset of the invisible credit population and we help make invisible credit a much smaller claim than it seems if the underlying mechanism only reaches the banked half of that group
The third thing and honestly the one I find hardest to model clearly is the adverse selection question we mentioned earlier in this article. When I see a lender advertise that it approves people that other lenders reject my first instinct used to be that they must have a better model. My read now is that they could instead be pricing adverse selection risk directly charging a rate that assumes the pool is skewed toward the riskier half of an unscored population and then advertising the approval rate as generosity.rather than a repriced bet. I can't always tell those two apart from the outside and I want to be honest: I don't have a clear test for that. The interest rate and fee structure are often the best clues since a rate created to absorb adverse selection tends to look expensive even when marketing presents it as inclusive
None of this makes me feel cynical about the space. Rental reporting and cash flow underwriting are real improvements not marketing disguised as improvements and the direction of travel here is a really good one for a lot of people who were excluded for reasons that had nothing to do with their actual reliability. But my honest conclusion is that we use alternative data as a description of a method not as proof of a result and the method deserves the same scrutiny as any other underwriting claim not a pass because the target population is sympathetic
The Bottom Line
Credit invisibility seems like an equity problem and is discussed as such but at heart it is a modeling problem. A scoring system based on payment data has nothing to calculate when the payment data does not exist and no file behaves mechanically exactly like a bad file even though it means something completely different. The math in the worked example is the entire mechanism in miniature: when a lender cannot distinguish a low-risk applicant from a high-risk one pricing for the average invites the safer half to walk away.and the riskiest half to stay so rejecting the entire group is often the rational not hostile response. Reporting rent entering cash flow data and products like Experian Boost really close that gap and deserve credit for it. They don't close it especially for people who are unbanked or paid in cash and expanding what counts as reportable data brings its own fair credit risk back into the picture. The honest bottom line is that this is an information problem before anything else and every solutionthat works it works by giving the model something new to look at not by changing what it does once it can see it