Prediction Without Explanation: How Much to Trust a Model That Is Right for Reasons No One Can State
A model that predicts returns and cannot say why is not a hypothetical case. It is the ordinary output of a modern research pipeline. Fit a gradient-boosted tree or a neural network to a few hundred features and a decade of cross-sectional returns, hold out the most recent years, and the machine will often rank securities better out of sample than any single relationship a researcher could write down. Ask what it has found and the honest answer is a list of inputs with weights attached, which is a description of the model rather than an explanation of the market. The uncomfortable question follows at once: if the forecasts keep being right, how much does the missing explanation matter?
There are two easy answers and both are wrong. The first is that prediction is all that matters, that markets pay for forecasts and not for stories, and that a demand for explanation is a form of nostalgia. The second is that an unexplained signal is a curve fit waiting to fail and should not be traded at all. This piece treats the question as a quantitative one. We ask what an explanation actually buys, why the data alone cannot replace it within a useful horizon, and how a research organization can turn the admission that it does not know why something works into a specific, smaller, and more closely watched allocation rather than a verdict.
The Trade-Off Is Real, and Smaller Than It Looks
Begin with the trade-off everyone assumes exists. Simple models are easy to explain and, the folk wisdom goes, less accurate; flexible models are more accurate and opaque. Both halves are partly true, and the relationship between them is not a straight line. Figure 1 places a set of stylized model classes on two axes. The horizontal axis is a rough interpretability score, running from a linear regression whose coefficients can be read directly, through tree ensembles whose partial dependences can be inspected one input at a time, to deep networks and sequence models in which the mapping from input to forecast has no compact description. The vertical axis is an out-of-sample information coefficient, the rank correlation between forecast and realized return, generated from the simple concave curve stated in the note.[1]
Note: Each model class is placed at a stylized interpretability score s in [0, 1] and given an out-of-sample IC of 0.020 + 0.028·√(1 − s) − 0.010·(1 − s)³, plus a class-specific offset of at most ±0.003 chosen to spread the points. The dashed line is the least-squares fit. Neither axis is measured from any dataset; the ordering of the classes is the point, not the values.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
Two features of the picture matter. The frontier rises as interpretability falls, but it flattens: in the stylized construction, most of the gain from flexibility arrives in the move from a single linear model to a well-regularized ensemble of shallow learners, whose behavior can still be interrogated. The move from there to fully opaque architectures adds less than the folk wisdom promises, and at the far end the additional flexibility begins to buy noise about as often as it buys signal. The cost of explanation, in other words, is real but modest, and it is paid mostly at one step of the ladder rather than spread evenly along it.
That changes what the question is about. If the price of an explanation were a large fraction of the available predictive power, a firm would have to choose between understanding and returns. If it is a modest fraction, the choice is between a slightly weaker model that can be reasoned about and a slightly stronger one that cannot, and the right answer depends on what reasoning is worth. The rest of this piece is an attempt to price it.
What an Explanation Is For
An explanation of a predictive relationship is not a courtesy to the reader. It is a statement of the conditions under which the relationship should hold: who is on the other side of the trade, what they are paid for or constrained by, and what would have to change for them to stop. A signal that works because index rebalancing produces price-insensitive flows on known dates arrives with its expiry conditions written in. A signal that works for reasons no one can state arrives with none, and the difference is invisible until the world changes.[2]
Figure 2 stylizes what happens at that point. Two signals with the same information coefficient are drawn through a regime shift at month zero: a change in market structure, a shift in the behavior of the participants who were supplying the return, a rule change. The explained signal's IC decays toward a floor with a long half-life, because the mechanism it depends on has been weakened rather than removed, and the researchers who own it know which part of the mechanism to watch. The unexplained signal decays with a short half-life toward a lower floor, because whatever it was exploiting is gone and nothing about the model says so. The parameters are chosen for exposition, but the shape follows from the argument: the explained signal degrades in a way its owners can anticipate, and the unexplained one degrades in a way they can only observe.
Note: Before month 0 both signals sit at an IC of 0.040. After the shift the explained signal follows 0.025 + 0.015·exp(−ln 2 · t / 24) and the unexplained signal 0.005 + 0.035·exp(−ln 2 · t / 4), each with a deterministic wobble of 0.003·sin(t / 1.7) + 0.002·cos(t / 3.9) added for realism (phase-shifted for the second series). The half-lives, floors, and threshold are chosen for exposition, not estimated from data.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
The obvious objection is that the decay will show up in live performance, and the model can be retrained or retired when it does. That is true, eventually. The question is how long eventually is, and here the arithmetic is unforgiving.
Why the Data Cannot Tell You in Time
A monthly information coefficient is a noisy quantity. Across a broad cross-section, the realized IC of a signal whose true value is a few hundredths swings by something like a tenth from month to month, and the standard error of its average over N months is that dispersion divided by the square root of N.[3] To conclude with ordinary statistical confidence that a signal's IC has fallen by half, one needs enough months for the expected drop to stand two standard errors clear of the noise. Figure 3 plots that requirement across a range of starting IC levels using the formula in the note. For a signal with an IC of 0.03, the kind of modest, broadly diversified relationship a machine-learned model typically produces, the requirement is measured in years, not quarters.
Note: Months required N = (z · σ / (f · IC))², with z = 1.96, a month-to-month standard deviation of realized IC of σ = 0.10, and f the fraction of the IC lost (0.5 for a halving, 1 for a complete loss). The σ is a round illustrative value; the curves are the formula, not an estimate from any dataset.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
This is the crux of the case for explanation, and it is quantitative rather than aesthetic. An unexplained signal cannot be retired on evidence within any horizon a portfolio can afford, because the evidence accumulates far more slowly than the losses do. It can only be retired on a rule set in advance, or on an argument. An explained signal supplies the argument: its owners can look at the mechanism directly, observe that the flow has stopped or the constraint has been lifted, and act long before the performance data could have told them anything. The explanation is a monitoring instrument with a much shorter lag than the returns themselves. That is what it is worth, and in a regime shift it is worth a great deal.
Five Questions in Place of an Explanation
None of this means an unexplained signal should be discarded. It means the missing explanation has to be replaced with something, and the something has to be more than a good backtest. Over time we have settled on five questions a signal must answer before it is trusted. Each can be asked of a model that cannot explain itself, and each recovers part of what an explanation would have provided. Figure 4 lays them out.
| Criterion | The question | Passes when | Fails when | What it stands in for |
|---|---|---|---|---|
| Stability | Does the relationship hold across periods, universes, and reasonable choices of hyperparameters? | The sign and rough size of the IC persist across decades, size groups, and regions, and do not hinge on one tuning choice. | The IC reverses in a subperiod or vanishes outside the largest names. | The knowledge that the mechanism is general rather than a feature of one sample. |
| Monotonicity | Do the model's partial dependences move in one direction, or in a direction someone can justify? | Forecasts respond smoothly to inputs, and the few non-monotone responses have a stated reason. | The response to an input bends for no reason, usually at a value seen in a handful of observations. | The check a person would make that the model uses an input the way the world does. |
| Ablation | If the inputs the model claims to rely on are removed and the model refit, does accuracy fall by the amount the claimed reliance implies? | Removing the headline inputs degrades the model materially; removing the rest does not. | The model performs as well without its headline inputs, so its own account of itself is wrong. | The causal part of an explanation: what the model is actually using. |
| Economic plausibility | Who is on the other side, and why would they accept the trade at a loss? | A counterparty with a constraint, mandate, or need for liquidity can be named, even conjecturally. | The only account of the return is that the model found it. | The mechanism itself, and with it the conditions under which it stops. |
| Capacity | Does the forecast survive the trading it would require at a size that would matter? | Predicted return exceeds modeled cost at the intended scale, and the signal is not concentrated in the least liquid names. | The signal is correct only where it cannot be traded. | The assurance that the return is a market return rather than a paper one. |
Note: Qualitative summary of the criteria described in the text. The entries describe what each test looks for, not thresholds; the thresholds depend on the horizon and breadth of the signal.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
Stability and monotonicity are the cheapest to check and the most often skipped. A relationship that holds in one decade and reverses in the next, or that holds among large-capitalization stocks and vanishes below them, is not a relationship but a coincidence with a long run. A model whose forecast rises with an input up to a point and then falls, for no reason anyone can supply, is usually fitting an artifact of the sample. Ablation asks whether the model's own account of what matters survives an intervention: remove the inputs it claims to depend on, refit, and see whether the loss in accuracy is what the claimed dependence implies.[4] Economic plausibility is the closest of the five to an explanation and the hardest to satisfy honestly; it asks not for a story but for a counterparty, someone whose constraints or incentives would lead them to take the other side of the trade at a loss. Capacity asks whether the signal survives the trading it would require, because a forecast that is only correct at a size too small to matter is a curiosity rather than a source of return.
A signal that passes all five is, for practical purposes, explained, even if the explanation was assembled from the outside rather than derived from the inside. A signal that passes stability and ablation but not plausibility is the genuinely hard case: it is real, it is not an artifact of the model, and no one knows why it works. That is the case the governance question is about.
Trust as a Budget, Not a Verdict
The mistake in both easy answers is to treat trust as binary. A model is deployed or it is not; an explanation is present or it is absent. In practice trust is a quantity, and the natural way to express it is as a risk budget and a review cadence that depend on how much of the explanatory gap remains open. Figure 5 sketches a stylized version of such a policy. Signals that pass all five questions receive the full budget their forecast strength would justify and are reviewed on the ordinary cycle. Signals with a conjectured but unverified mechanism receive a fraction of that and a faster review. Signals with no mechanism at all receive a smaller fraction still, a review cadence short enough to catch the decay in Figure 2 before it matters, and an automatic retirement rule keyed to a drawdown or a stability failure rather than to anyone's judgment. Signals that fail stability are not deployed, however good the backtest.
1×
Tier 1: explained
All five criteria pass and the mechanism is written down; full budget; reviewed on the ordinary quarterly cycle
½×
Tier 2: mechanism conjectured
Stable, monotone, and ablation-consistent; counterparty proposed but unverified; half budget; reviewed monthly
¼×
Tier 3: unexplained
Stable and ablation-consistent only; quarter budget; reviewed weekly; automatic retirement on a preset drawdown or a stability failure
0×
Fails stability
Not deployed, whatever the backtest shows
Note: A hypothetical policy drawn to illustrate the principle in the text. The multipliers scale the risk budget a signal's forecast strength alone would justify; the tiers, review cadences, and retirement rules are illustrative and are not a statement of any actual limit.
Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.
The multipliers in the figure are illustrative and the tiers could be drawn differently. The principle is not. An unexplained signal is not worthless, but it is worth less than an explained signal of the same strength, by an amount that reflects how much longer it will take to learn that it has stopped working. Sizing it accordingly is not caution for its own sake. It is the expected-value calculation that the missing explanation forces, done in advance rather than after the fact.
What This Means for How We Work
Three practices follow. The first is to treat explanation as a research output in its own right, with time budgeted for it after a model is built rather than only before: ablations, stability maps across periods and universes, and a written statement of the conjectured counterparty, even when the conjecture is weak. The second is to design the monitoring of an unexplained signal around the arithmetic of Figure 3, which means watching the inputs and the market structure it depends on rather than its returns, because the returns will be the last thing to tell you. The third is to keep the question open. Some of the relationships a model finds without explanation acquire one later, when someone notices which flow or constraint they were tracking, and the work of finding out is often where the next explained signal comes from.
The uncomfortable question at the start has a plain answer. A model that predicts and cannot explain should be trusted in proportion to how much of the explanation can be reconstructed from the outside, and for no longer than it takes the market to change. That is less trust than a good backtest invites and more than a purist would allow. It is also, as far as we can tell, the amount the arithmetic supports.
- [1]The interpretability score is an ordering rather than a measurement; no agreed scale exists. We follow the convention in the interpretability literature that a model is more interpretable the more of its behavior can be reconstructed from a description a person can hold in mind (Lipton, 2018; Rudin, 2019). The information coefficient follows Grinold and Kahn (2000): the correlation between a forecast and the subsequent realized return.
- [2]The idea that a predictive relationship is conditional on a population of market participants that itself evolves is developed in Lo's adaptive markets hypothesis (2004). A regime shift in this piece's sense is a change in that population or in its constraints.
- [3]For a true monthly IC of 0.03 and a month-to-month standard deviation of realized IC of 0.10, the two-standard-error band around a 24-month average is roughly ±0.04, wider than the IC itself. The same dispersion that limits inference is why Grinold and Kahn's fundamental law rewards breadth: many small, noisy forecasts add up even when no single one can be verified quickly.
- [4]Ablation, in which an input is removed and the model refit, is distinct from importance measures computed on a fixed model, which can credit inputs that are merely correlated with the ones doing the work. The cost of ablation is the refit; the benefit is that it tests an intervention rather than an association.
Interested in related insights?
Regime Change: Are Market States Real, and Can They Be Recognized Before They End?
The Backtest That Never Happened: How Historical Simulations Learn What No One Could Have Known
Enjoyed this piece?
This document is provided for informational purposes only and does not constitute investment advice or an offer to sell (or the solicitation of an offer to buy) any security, investment product, or service.
The views expressed are those of OAK ST LLC as of the date of the document, are subject to change without notice, and may not reflect the criteria used by OAK ST LLC to evaluate investments. Figures described as illustrative, stylized, or simulated are hypothetical constructions prepared for exposition; they do not depict the results of any OAK ST LLC strategy, portfolio, or account, and no representation is made that any account will or is likely to achieve results similar to those shown. Historical market trends are not reliable indicators of future market behavior.
Information obtained from third-party sources is believed to be reliable but has not been independently verified, and OAK ST LLC does not guarantee its accuracy or completeness. Nothing in this document is a recommendation to buy, sell, or hold any instrument.
This document may not be reproduced or distributed without the prior written authorization of OAK ST LLC. The Terms of Use and the Important Legal and Regulatory Disclosures govern its use. Copyright © 2026 OAK ST LLC. All rights reserved.