The Information Budget: How Many Independent Things Do Thousands of Datasets Actually Say?

The catalog of data that can be bought is now longer than any research group can read. Card transactions, web traffic, satellite imagery, job postings, product prices, shipping manifests, the text of every filing and every news wire: each arrives as a separate product with a separate contract and a separate claim to be predictive. It is natural to count datasets and to treat the count as a measure of informational advantage. It is also wrong. A dataset is a way of observing the economy, and most ways of observing the economy see the same few things. Two panels of consumer spending are not two sources of information about consumption; they are two noisy views of one.

This piece asks how much genuinely independent information survives once thousands of datasets are netted against the common economic exposures they share and against each other. We work in a stylized universe rather than a real one, because the question is about structure, not about any particular product. The answer, in that stylized world, is that the count of independent dimensions is smaller than the catalog by roughly two orders of magnitude, that the price of an independent dimension varies by close to an order of magnitude across dataset families, and that this count, rather than the length of the catalog, is the budget a systematic research process actually operates within.

Breadth Is Counted in Independent Bets

The reason the count matters comes from the arithmetic of active management. The value of a forecasting process depends on the skill of each forecast and on the number of independent forecasts it makes.[1] The second term is breadth, and breadth is not the number of datasets or the number of signals. It is the number of bets that are not already implied by the other bets. Ten datasets that all track consumer spending, each with a different lag and a different sampling error, add skill to a single consumption forecast. They do not add breadth. Once that forecast is made, the eleventh spending panel is worth only the marginal reduction in its noise.

The structure of a dataset universe can be read from its correlation matrix. If every dataset were independent, the eigenvalues of that matrix would be roughly equal, and the cumulative variance explained by the leading principal components would rise along a diagonal (a straight line on linear axes; on the logarithmic component axis used in Figure 1 it bows upward toward the right). If a handful of common exposures dominated, the first few components would explain most of the variance and the curve would rise steeply, then flatten. Figure 1 draws that curve for a stylized universe of two thousand dataset-level series built from twelve common factors and a long tail, and draws it again after the twelve factors have been regressed out of every series.

Figure 1:  Cumulative Variance Explained by Principal Components, Before and After Removing Common FactorsStylized universe of 2,000 dataset-level series; log scale on the component axis
0%25%50%75%100%1310301003001000200012 common factors≈ 50 structured dimensionsCumulative variance explainedNumber of principal components
Raw universe (share of total variance)After removing 12 common factors (share of residual variance)Equal-variance benchmark (k ÷ N)

Note: Raw spectrum: twelve common factors with eigenvalues 0.18 × 0.75^(k−1), k = 1…12 (about 70% of total variance), followed by a residual spectrum μ_j = 0.012 × exp(−j/40) + 0.00035, j = 1…1,988, scaled to the remaining 30%. The residual curve is the same μ_j renormalized to sum to one, as it would be after regressing every series on the twelve factors. The benchmark is k ÷ N. All parameters are chosen for exposition.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

Two things stand out. The raw curve reaches roughly seventy percent of the variance by the twelfth component, which is what a universe dominated by the market, by sector, by the consumer cycle, and by a few macroeconomic conditions looks like. Everything a dataset says about those exposures is already said by prices, by earnings, and by the other datasets. The residual curve is far flatter, but it is not the diagonal. It bends above the pure-noise benchmark for its first several dozen components, after which it rises only with the noise floor and the benchmark slowly catches up with it. That bend is the part of the universe that is both independent of the common factors and shared across several datasets in a structured way: the candidate independent dimensions. In this construction there are on the order of fifty of them, out of two thousand series.[2]

Principal components count dimensions of variance, not dimensions of information. A direction that explains a great deal of residual variance may forecast nothing; a direction that forecasts something useful may be small. The scree is therefore an upper bound on the count of independent things a universe can say, and the true count is found only by testing each candidate direction out of sample against the forecasts the firm already makes. That is slow, and its slowness is the first constraint on the information budget.

Two Layers of Redundancy

The scree collapses every dataset into one picture. The more useful decomposition is by family, because redundancy has two layers and they weigh differently on different kinds of data. The first layer is the common economic exposure: the share of a dataset's variance explained by the factors every dataset sees. The second is the within-family layer: the share of what remains that is shared with other datasets of the same kind, because three card panels sample the same consumers, three geolocation providers count the same visits, and every weather feed reports the same weather.

Figure 2 stacks the two layers for ten stylized families. The residual slice, the part of a dataset's variance that is specific to that dataset after both layers are removed, is what a firm actually pays for when it buys a second or third source in a family it already holds. The parameters are chosen to illustrate the mechanism, not estimated from any real data, but the ordering follows from the nature of what each family measures.

Figure 2:  Where a Dataset's Variance Goes: Common Factors, Family, and Dataset-Specific ResidualTen stylized dataset families; share of variance
0%25%50%75%100%Card panelsWeb trafficGeolocationSatelliteJob postingsProduct pricingNews sentimentSupply chainWeatherFilings text
Common economic factorsShared within the familyDataset-specific residual

Note: For each family, the common-factor share is an illustrative R² from regressing a typical series on the twelve common factors of Figure 1; the within-family share is (1 − R²) × ρ, where ρ is the illustrative fraction of the remainder shared with other datasets of the same kind; the dataset-specific residual is (1 − R²) × (1 − ρ). R² ranges from 0.15 (weather) to 0.74 (news sentiment) and ρ from 0.40 (filings text) to 0.80 (weather); both are chosen to illustrate the ordering, not estimated.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

The pattern repays study. Sentiment extracted from news and social text loads heavily on the common factors, because most of what moves sentiment is what moves the market, and the several providers who extract it apply similar methods to the same text; the dataset-specific residual is small. Weather sits at the opposite extreme on the first layer, since weather does not care about the business cycle, and at the same extreme on the second, since every provider is measuring the same atmosphere; a second weather dataset is almost entirely redundant with the first. Satellite imagery and filings text carry the largest dataset-specific residuals in this construction, because what is extracted from an image or a document depends more on the extraction than on the underlying object, and different extractions genuinely disagree.

None of this says which families are valuable. A small residual can still be a valuable residual if it forecasts well, and a large one can be pure measurement noise. It says where the marginal dataset is likely to add a new dimension and where it is likely to add a noisier copy of one the firm already holds. In a research process that must choose what to evaluate, that is most of what is needed.

The Price of an Independent Bit

Data is priced per dataset. Information arrives per independent dimension. The gap between those two units is where the economics of alternative data lives. Figure 3 works the arithmetic for the same ten stylized families: how many datasets are screened, how many survive screening for coverage, history, and point-in-time integrity, how many independent dimensions the survivors add once they are netted against the common factors and against each other, and what that implies for the cost of a dimension relative to the cost of a dataset. We use the word bit loosely here, to mean one independent dimension of predictive information rather than a unit of entropy.[3]

Figure 3:  A Stylized Cost per Independent Bit, by Dataset FamilyCounts and cost indices for the ten families of Figure 2
Dataset familyScreenedRetainedIndependent dimensionsCost per dataset (index)Cost per independent dimension (index)
Card and receipt panels24082100400
Web and app traffic38012360240
Geolocation and foot traffic2006280240
Satellite imagery9053120200
Job postings and hiring1605240100
Product pricing and catalogs3201043075
News and social sentiment4806150300
Shipping and supply chain1906270210
Weather and physical110311545
Filings and regulatory text230942556
All families24007024170

Note: Screened, retained, and independent-dimension counts describe the stylized universe of Figure 2. Cost per dataset is an arbitrary index (weather and physical = 15) with no currency attached; cost per independent dimension = retained × cost per dataset ÷ independent dimensions. The final row totals the universe, with the total cost per dimension computed as total retained cost ÷ total dimensions.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

The cost per independent dimension is not the cost per dataset. Families whose datasets are highly redundant, card panels and sentiment feeds in this construction, produce the most expensive dimensions, because the firm pays for several sources to obtain one or two directions that are new. Families whose datasets are cheap and structurally diverse, product catalogs and filings text, produce the cheapest dimensions after weather, whose datasets are close to free, even though no single dataset in either family looks impressive. The spread across families in the table is nearly an order of magnitude, and it does not follow the spread in headline prices: the most expensive dataset in the table, satellite imagery, buys a mid-priced dimension, while a mid-priced sentiment feed buys the second most expensive one.

There is a reason the expensive data tends to be the redundant data. A dataset commands a high price because many buyers want it, and many buyers want it because it measures something obviously important. Something obviously important is, almost by definition, something the market already tracks through other channels. The residual is small because the exposure is large. The economics of a data catalog therefore have the same shape as the economics of a crowded trade: the sources that are easiest to justify buying are the ones whose content is most likely to be already in the price.

The Budget

Put the pieces together and the size of the information budget becomes visible. Figure 4 totals the stylized universe: the datasets screened, the datasets retained, the independent dimensions they add, and the ratio between the first and the last.

Figure 4:  The Information Budget of a Stylized Data UniverseTotals of the ten families in Figure 3
  • 2,400

    Datasets screened

    ten stylized families

  • 70

    Retained after screening

    ≈ 3% of those screened

  • 24

    Independent dimensions

    after netting against common factors and each other

  • 100 : 1

    Datasets screened per independent dimension

    2,400 ÷ 24

Note: Totals of the stylized universe in Figure 3. Retention and independent-dimension counts are constructions chosen to illustrate the two layers of redundancy, not counts from any real screening process.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

One dimension per hundred datasets screened is not a pessimistic number; in a stylized world with two layers of redundancy it is the natural one. What it implies is that the size of the catalog is not the measure of a data program. The measure is the count of independent dimensions the program has validated, the rate at which it can add to that count, and the cost of each addition. All three are bounded, and the bounds tighten as the inventory grows, because every candidate must be evaluated against everything already held. The hundredth dataset has to be tested for novelty against ninety-nine; the thousandth against nine hundred and ninety-nine. The cost of admission rises with the size of the room.

The same arithmetic constrains the models built on top of the data. A learning algorithm given thousands of features does not receive thousands of dimensions; it receives a few dozen, wrapped in noise, with the noise dressed up as structure. The number of parameters it can support without fitting that noise is set by the independent information in the inputs, not by their count, and the multiple-testing problem that afflicts signal research afflicts feature selection in exactly the same way.[4] This is the practical reason we regularize by the information budget rather than by the feature count: a model is allowed as much complexity as the independent dimensions can pay for, and no more. Adding features that add no dimensions does not make the model smarter. It makes the false-discovery arithmetic worse.

What This Means for How We Work

Four practices follow. The first is to residualize before evaluating: a candidate dataset is judged on what it says after the common factors and the existing inventory have been removed, never on its raw correlation with returns, because the raw correlation is mostly a restatement of exposures already held. The second is to price per dimension: the question put to a data budget is not what a dataset costs but what an independent dimension in its family costs, and the answer is allowed to reject an inexpensive dataset in a redundant family and accept an expensive one in a diverse family. The third is to retire: a dataset whose residual has been absorbed by a newer or cleaner source is a cost without a dimension, and the inventory is pruned as deliberately as it is grown. The fourth is to keep the count explicit. A research group that knows how many independent things its data can say, and how fast that number is changing, can allocate its attention. One that counts datasets cannot.

The catalog will keep growing. The number of independent things the economy does will not grow with it. The distance between those two facts is the information budget, and a systematic research process is, in the end, an organization for spending it well.


  1. [1]The relationship is the fundamental law of active management of Grinold and Kahn: the information ratio of a process is approximately its information coefficient multiplied by the square root of its breadth, where breadth is the number of independent forecasts per year. Datasets that load on the same exposure contribute to one forecast, and therefore to the coefficient, not to the breadth.
  2. [2]The diagonal is a loose benchmark. Even a universe of pure noise produces a spread of eigenvalues in a finite sample, with the largest well above the average, and the appropriate comparison is the eigenvalue distribution of a random matrix of the same dimensions, described by Marchenko and Pastur. In practice the residual spectrum is compared against that distribution rather than against the diagonal, and the count of components that clear it is smaller still.
  3. [3]Strictly, the information a dataset carries about future returns is a mutual information, measured in bits, and per observation it is very small for any dataset. The loose usage in the text, one bit for one independent dimension of predictive content, is a convenience for talking about budgets.
  4. [4]Harvey, Liu, and Zhu make the case that with hundreds of published factors the threshold for accepting a new one must rise; Bailey and López de Prado make the corresponding point for backtests. The argument transfers directly to datasets and to features: the more candidates that have been examined, the higher the bar each must clear.

Interested in related insights?

The Cost of Knowing Too Much: Does a Richer Model Predict Better, or Only Fit Better?

When Everyone Sees the Same Signal: How a Good Prediction Becomes a Bad Trade

Enjoyed this piece?

Share your thoughts!

This document is provided for informational purposes only and does not constitute investment advice or an offer to sell (or the solicitation of an offer to buy) any security, investment product, or service.

The views expressed are those of OAK ST LLC as of the date of the document, are subject to change without notice, and may not reflect the criteria used by OAK ST LLC to evaluate investments. Figures described as illustrative, stylized, or simulated are hypothetical constructions prepared for exposition; they do not depict the results of any OAK ST LLC strategy, portfolio, or account, and no representation is made that any account will or is likely to achieve results similar to those shown. Historical market trends are not reliable indicators of future market behavior.

Information obtained from third-party sources is believed to be reliable but has not been independently verified, and OAK ST LLC does not guarantee its accuracy or completeness. Nothing in this document is a recommendation to buy, sell, or hold any instrument.

This document may not be reproduced or distributed without the prior written authorization of OAK ST LLC. The Terms of Use and the Important Legal and Regulatory Disclosures govern its use. Copyright © 2026 OAK ST LLC. All rights reserved.