Counting Cars in Parking Lots Became a Serious Investment Input
Alternative data means information about a company that does not come from the company. It moved from novelty to standard practice, and the returns to using it decayed exactly as theory predicts.
The Basic Idea
Companies report results quarterly, weeks after the period ends. Between those reports, investors are guessing. Alternative data is any information that narrows that guess without waiting for the filing, and which does not come from the company itself.
The categories are broad. Satellite and aerial imagery of car parks, shipping terminals, and construction sites. Aggregated and anonymised card transaction panels. Web traffic, app downloads, and job postings. Shipping manifests and customs records. Geolocation footfall data.
Why It Works at All
Each of these is a proxy, meaning something correlated with the number you care about but not the number itself. Cars in a car park correlate with store visits, which correlate with sales. Job postings for warehouse staff correlate with expected volume.
The value is in timeliness rather than accuracy. A rough estimate available six weeks before the official figure can be worth far more than a precise figure everyone receives simultaneously.
The edge was never in knowing the number exactly. It was in knowing roughly which direction it moved before anyone else had any information at all.
The Problems Are Substantial
The difficulties are larger than the popular account suggests, and they are why most of these datasets disappoint.
| Problem | Why it undermines the signal |
|---|---|
| Coverage bias | A card panel skews to certain demographics |
| Short history | Too few quarters to validate reliably |
| Changing composition | The panel itself shifts, mimicking real change |
| Proxy breakdown | Online orders detach store visits from sales |
The coverage issue is the most common failure. A transaction panel covering a few percent of consumers, concentrated in particular card issuers and income brackets, may track a company overall sales well for years and then diverge when the customer mix changes. Nothing announces that the relationship has broken.
Short history compounds it. Validating that a signal predicts results requires many observations, and a dataset with three years of history offers roughly twelve quarterly data points, which is nowhere near enough to distinguish a real relationship from a coincidence.
The Decay
The economics of this industry follow a predictable arc. Early users of a dataset earn real excess returns. The data vendor, seeing demand, sells to more clients because incremental sales cost almost nothing. As adoption spreads, the information is impounded into prices faster, and the advantage compresses toward zero.
What remains is not an edge but a requirement. Once a dataset is widely used, not having it means being systematically late, so firms pay for it to avoid a disadvantage rather than to gain one. That is a meaningful and expensive distinction.
The Legal and Ethical Boundary
The constraint that shapes the whole industry is material non public information. Trading on confidential information obtained in breach of a duty is illegal, and alternative data has to stay clearly on the other side of that line.
The general principle is that independently observed information about a company is legitimate, while information originating inside the company or its partners in breach of an obligation is not. Satellite imagery of a public car park is observation. A dataset assembled from a supplier internal systems in violation of its contracts is a serious problem, and the buyer bears risk for it.
Consumer privacy adds a second constraint. Data derived from individuals requires proper anonymisation and consent, and regulation here has tightened considerably.
What Good Practice Looks Like
Serious users treat these datasets as one input rather than a signal to trade directly. They test whether the relationship to reported results holds across many quarters and many companies, they understand the collection methodology well enough to know when it would break, and they monitor for the composition shifts that quietly invalidate a history.
The failure mode is fitting a story to a short history and trusting it, which is the same mistake as any overfitted model, arriving through a more expensive channel.
The Bottom Line
Alternative data is a legitimate attempt to observe business activity directly rather than waiting to be told about it. The information is real, the proxies are imperfect and occasionally break without warning, and the advantage erodes as adoption grows. It has largely become the cost of staying current rather than a way of getting ahead.