The Ensemble Advantage: Why Combinations of Imperfect Models Beat Their Best Member, and When Adding One More Hurts
A committee of imperfect models can outperform its most accurate member, and in forecasting financial returns it usually does. The result is old, it is robust, and it still surprises people who meet it for the first time. If model A predicts returns better than model B, the intuition runs, then mixing B into A can only dilute it. The intuition is wrong because it treats accuracy as the only thing a model contributes. A forecast is a signal wrapped in error, and when two models make different errors, averaging them cancels part of the error while preserving the signal. What matters for the combination is not how good each model is on its own but how much of each model's error the other does not share.
This piece sets out the arithmetic of that effect in the simplest useful form, shows what it implies for the models a systematic investor should be looking for, and then turns to the less discussed half of the question: when adding another model makes the system worse. Both halves follow from the same few quantities, the strength of each model, the correlation among their errors, and the cost each model imposes on the whole. The conclusion is that the ensemble advantage is real but bounded, that it is earned by finding models that disagree in the right way rather than by accumulating models, and that a research process needs a rule for retiring models as much as it needs one for admitting them.
Why an Average Beats Its Best Member
Take N forecasting models, each producing a standardized forecast of next period's return for every security in a universe, and give each the same modest skill: an information coefficient, the correlation between its forecasts and the returns that follow, of 0.03. Suppose the forecasts of any two models are correlated with one another at ρ. The equal-weighted average of the N forecasts has the same covariance with returns as any one model, because covariances average. Its variance is smaller: the average of N unit-variance forecasts has variance (1 + (N − 1)ρ) / N. Dividing the covariance by the square root of that variance gives the ensemble's IC, which is the single model's IC multiplied by the square root of N / (1 + (N − 1)ρ).[1]
Figure 1 draws that quantity for average correlations between 0.1 and 0.9. Two features stand out. The gains diminish quickly: most of what an ensemble will ever deliver is delivered by the first handful of members. And the gains have a ceiling. As N grows without limit the multiplier approaches 1 / √ρ, so models whose forecasts are correlated at 0.9 can never combine into anything much better than one of them, while models correlated at 0.1 could, in principle, combine into something three times as strong. The ceiling is set entirely by correlation. No number of additional models moves it.
Note: Each curve is 0.03 × √(N / (1 + (N − 1)ρ)), the IC of an equal-weighted average of N unit-variance forecasts that each have IC 0.03 and share the pairwise correlation ρ given in the legend. The limit as N grows is 0.03 / √ρ.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
The comparison that matters in practice is not between the ensemble and an average member but between the ensemble and its best member. In the stylized model the members are identical, so the ensemble beats the best by construction. Make them unequal, with ICs spread from 0.02 to 0.04 around the same mean, and the equal-weighted ensemble still beats its best single member for any correlation below roughly 0.45, because the diversification gain outweighs the drag of the weaker members. This is the ensemble advantage in its plain form: the best model in a set is rarely as good as the set, and identifying which model is best is itself a noisy exercise that the ensemble sidesteps.
Correlation Comes From Inputs, Not Algorithms
If correlation sets the ceiling, the question becomes where correlation comes from. The answer, in our experience, is inputs. Two models built on the same data make the same errors, however different their architectures, because the errors live in the data: the stale accounting figure, the corporate action mishandled in the price history, the news item everyone saw at the same moment. A gradient-boosted tree fit to fundamental ratios is not a different view of the world from a linear value model. It is a different summary of the same view, and its errors are largely the value model's errors rearranged.
Figure 2 illustrates the pattern with a stylized correlation matrix for eight model families. The matrix is generated by a simple latent-factor construction described in the note: each model's forecast error loads on four underlying information sources, fundamentals, price history, order flow, and text, plus a private component, and the correlation between two models is the overlap of their loadings.[2] The construction is invented, but the structure it produces is the one we expect to see: blocks of high correlation among models that share a data source, near-zero correlation across sources, and the two machine-learning models sitting inside the block of whatever they were trained on.
Note: Each model's error loads on four latent sources (fundamentals, price history, order flow, text) plus a private component; the correlation between two models is the inner product of their loadings divided by the product of their total standard deviations. Loadings and private volatility: Value 0.9, 0.1, 0, 0.1 (0.5); Quality 0.8, 0, 0, 0.2 (0.6); Momentum 0.1, 0.9, 0.1, 0 (0.5); Reversal 0, −0.4, 0.6, 0 (0.6); Order flow 0, 0.3, 0.9, 0 (0.5); News text 0.2, 0.2, 0.1, 0.9 (0.5); Trees on fundamentals 0.9, 0.2, 0, 0.1 (0.35); Net on price history 0.1, 0.9, 0.2, 0 (0.35).
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
Two cells deserve attention. The tree model on fundamentals is correlated with the linear value model at roughly 0.8 in this construction, which means that adding it to an ensemble already containing the value model buys very little. The short-horizon reversal model is mildly negatively correlated with the momentum model, because the two respond to the same price history in opposite directions, and a model whose errors run against the ensemble's is worth more than one whose errors are merely unrelated. The practical lesson is that diversity is bought with new data and new horizons far more reliably than with new estimators. A more powerful learner on the same inputs raises the ensemble's ceiling by a little; a mediocre model on inputs no other member sees raises it by a lot.
A Weaker Model Can Be Worth More Than a Stronger One
The consequence for how candidates should be judged is uncomfortable. The natural way to evaluate a new model is by its stand-alone IC, and that number is close to irrelevant. What determines whether the candidate improves the ensemble is its IC relative to the incumbents together with its correlation with them, and of the two, correlation dominates.
Figure 3 makes the trade-off explicit. The incumbent ensemble contains five models of the kind used in Figure 1, correlated with one another at 0.4. A candidate is added at equal weight, and the chart shows the change in the ensemble's IC as a function of the candidate's correlation with each incumbent, for three candidates: one weaker than the incumbents, one equal to them, and one half again as strong. In the stylized model a candidate with two-thirds the skill of the incumbents improves the ensemble as long as its correlation with them is below roughly 0.3. A candidate half again as strong as the incumbents makes the ensemble worse once its correlation with them exceeds roughly 0.8. Between those points a weaker, different model beats a stronger, similar one.
Note: The six-model average has covariance (5 × 0.03 + IC_c) / 6 with returns and variance (5 + 20 × 0.4 + 1 + 10 × ρ_c) / 36, where IC_c is the candidate's IC and ρ_c its correlation with each incumbent; the chart shows the resulting IC relative to the five-model ensemble's IC of 0.03 × √(5 / 2.6), minus one.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
The chart also explains why ensembles are so often built badly. Candidates are generated by the same researchers, from the same data, using the same toolkit, and are then selected on stand-alone performance. That process is a machine for producing strong, similar models, which is exactly the region of the chart in which adding a model hurts. A research process that wants the ensemble advantage has to evaluate candidates against the ensemble rather than against zero, and it has to reward the researcher who brings a new data source with a modest IC at least as much as the one who squeezes a higher IC out of a familiar one.
When the Next Model Makes the System Worse
Figure 3 was drawn without costs, and costs are where the second half of the question lives. Every model added to a production ensemble imposes three kinds of charge. It adds turnover: a fast model mixed into slow ones lowers the persistence of the combined forecast, so the portfolio trades more often to follow a signal that is only marginally better. It adds estimation error: each model brings a weight to estimate, and the covariance-aware schemes discussed below bring a row and column of correlations too, all estimated from finite history and all leaking realized IC. And it adds operational burden: another data feed to maintain, another fit to monitor, another way for the system to fail quietly while its forecasts continue to arrive.
Figure 4 combines the diminishing gross gains of Figure 1 with a flat frictional charge per model, expressed in units of a single model's IC. The gross contribution of the k-th model falls with k, as it must; the charge does not. In the stylized version the net contribution turns negative at the sixth model, and everything added after that point makes the system worse while appearing, on a gross basis, to make it better. The exact crossing depends on parameters nobody knows precisely. The shape does not, and the shape is the point: there is always a k beyond which the next model is a liability, and an ensemble that has never retired a model has almost certainly passed it.
Note: Gross bars are [IC(k) − IC(k − 1)] / 0.03, where IC(k) = 0.03 × √(k / (1 + 0.4 (k − 1))) is the equal-weighted ensemble IC of k models from Figure 1 with ρ = 0.4. Net bars subtract a constant charge of 3% of a single model's IC per model, standing in for the turnover, estimation error, and operating burden each additional model adds.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
This is also the answer to a question we are asked often, which is why a firm with the ability to build hundreds of models would choose to run far fewer. The arithmetic of combination rewards the first few uncorrelated models generously and everything after that barely at all, while the cost of each model is roughly constant. A large ensemble of similar models is not more robust than a small ensemble of different ones. It is more expensive, harder to understand, and no better at forecasting.
How the Members Are Combined
Given a set of members, the weights can be chosen in several ways, and the choice matters less than people expect and in a direction they do not. Figure 5 summarizes the four schemes we consider most often. Equal weighting estimates nothing beyond the membership itself. Weighting by IC tilts toward the stronger members but ignores correlation. The inverse-covariance scheme, the forecasting analogue of a mean-variance portfolio, rewards diversity and penalizes redundancy but requires a full correlation matrix and behaves badly when two members are near-duplicates. Stacking hands the weights to a second-stage model that can learn conditional strengths, which member works in which environment, at the price of a further layer of fitting that must be validated strictly out of sample.
| Scheme | Weights | What must be estimated | What it rewards | How it fails |
|---|---|---|---|---|
| Equal weight | 1 / N for every member | Nothing beyond the membership | Robustness; hard to beat out of sample when skills are similar | Wastes a clearly stronger member; a broken member keeps full weight |
| IC-weighted | Proportional to each member's estimated IC | N information coefficients | Skill, as measured in sample | ICs are noisy, so it overweights whichever member was lucky; ignores redundancy |
| Inverse-covariance | Proportional to the inverse correlation matrix times the IC vector | N ICs and N (N − 1) / 2 correlations | Diversity; penalizes near-duplicates | Unstable when members are similar; can assign large weights of opposite sign |
| Stacked | Learned by a second-stage model from the members' forecasts | A further model, often nonlinear, with its own parameters | Conditional strength: which member works when | Leaks in-sample fit unless trained strictly out of fold; hardest to audit |
Note: Qualitative summary of the mechanisms described in the text. The schemes are listed roughly in order of how much they estimate, which is also, in our experience, the reverse of how often they win out of sample.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
The ordering of that table runs the wrong way for anyone who enjoys sophistication. Equal weights are hard to beat because the quantity the cleverer schemes estimate, the difference in skill and redundancy among members, is small relative to the noise in its estimate.[3] A scheme that tilts toward the member with the highest in-sample IC is, more often than not, tilting toward the member that was luckiest. Our practice is to treat equal weight as the null, to deviate from it only where the evidence for a difference among members is strong and stable, and to prefer removing a weak or redundant member outright to down-weighting it, since removal is a decision that can be audited and a fractional weight is not.
What This Means for How We Work
Four practices follow. The first is to score every candidate model by its marginal contribution to the ensemble it would join, net of the turnover and estimation cost it would add, rather than by its stand-alone performance; a candidate is a proposal to change a system, not a contestant in a competition. The second is to spend research effort on inputs before estimators: a new source of information with a modest IC is usually worth more than a better learner on a familiar one, because it lowers the average correlation and raises the ceiling on what the whole can achieve. The third is to hold membership to a budget, reviewing the ensemble regularly and retiring members whose contribution has decayed or been duplicated, so that the count of models tracks the count of independent views rather than the count of ideas anyone has had. The fourth is to keep the combination rule simple enough to explain in a sentence, because a rule that cannot be explained cannot be diagnosed when it fails.
None of this is a case against building many models. It is a case for combining few, chosen for how they differ, and for treating the ensemble rather than any of its members as the object of research. The advantage is real. It belongs to whoever is most disciplined about what they add and most willing to take away.
- [1]With unit-variance forecasts, each correlated at IC with the standardized return and at ρ with one another, the average has covariance IC with the return and variance (1 + (N − 1)ρ) / N, so its correlation with the return is IC × √(N / (1 + (N − 1)ρ)). The formula assumes equal ICs and a single common correlation; with unequal values the same logic applies to the averages and the qualitative conclusions do not change. It is the forecasting counterpart of the diversification arithmetic in Markowitz (1952) and of the breadth term in Grinold and Kahn's fundamental law of active management.
- [2]For weak forecasts the correlation between two models' forecasts and the correlation between their errors are nearly identical, since the shared signal component is of order IC² and negligible, and we use the two interchangeably. The latent-factor construction behind Figure 2 is the device used in factor risk models, applied to forecast errors rather than returns. The decomposition of an ensemble's error into the average member's error less the members' disagreement is due to Krogh and Vedelsby (1995); the general case for combining forecasts goes back to Bates and Granger (1969).
- [3]The finding that simple averages are hard to beat out of sample is old enough to have a name, the forecast combination puzzle; Timmermann (2006) surveys it. It is closely related to the observation that naive 1 / N portfolios are hard to beat with estimated mean-variance weights (DeMiguel, Garlappi, and Uppal, 2009), and for the same reason: the inputs to the optimization are estimated with errors large enough to swamp the gains from optimizing.
Interested in related insights?
Prediction Without Explanation: How Much to Trust a Model That Is Right for Reasons No One Can State
A Million Small Bets: Breadth, Skill, and the Fine Print of the Fundamental Law
Enjoyed this piece?
This document is provided for informational purposes only and does not constitute investment advice or an offer to sell (or the solicitation of an offer to buy) any security, investment product, or service.
The views expressed are those of OAK ST LLC as of the date of the document, are subject to change without notice, and may not reflect the criteria used by OAK ST LLC to evaluate investments. Figures described as illustrative, stylized, or simulated are hypothetical constructions prepared for exposition; they do not depict the results of any OAK ST LLC strategy, portfolio, or account, and no representation is made that any account will or is likely to achieve results similar to those shown. Historical market trends are not reliable indicators of future market behavior.
Information obtained from third-party sources is believed to be reliable but has not been independently verified, and OAK ST LLC does not guarantee its accuracy or completeness. Nothing in this document is a recommendation to buy, sell, or hold any instrument.
This document may not be reproduced or distributed without the prior written authorization of OAK ST LLC. The Terms of Use and the Important Legal and Regulatory Disclosures govern its use. Copyright © 2026 OAK ST LLC. All rights reserved.