Regime Change: Are Market States Real, and Can They Be Recognized Before They End?

Every market commentary is written in the language of regimes. Risk is on or off; volatility is in a low regime or a high one; rates have entered a new regime and the old playbook no longer applies. The word does real work. It says that the rules have changed, not merely the numbers. But its usual use is retrospective: a regime is named after it has ended, when all of the data that define it are in hand. The question a systematic investor has to ask is different. Do regimes exist as something other than a story told afterward, and if they do, can they be recognized while they are still in progress?

This piece takes the question up with the simplest tool that makes it precise, a two-state hidden Markov model of daily returns. We use it to say what a regime is, to separate what can be known in real time from what can be known only in hindsight, to measure how long recognition takes, and to show the trade-off, unavoidable in any detector, between false alarms and missed regimes. The conclusion is that regimes are real in a specific and limited sense; that they can usually be recognized early enough to matter; and that the price of recognizing them is a stream of false alarms that has to be budgeted for rather than wished away.

A Regime Is a Hidden State, Not a Label

A regime, in the sense that makes the word useful, is a state of the world that is not observed directly but that governs the distribution of what is observed. The state is hidden: no series in the market data is labeled with it. It is persistent: today's state is more likely than not to be tomorrow's. And it matters: the distribution of returns, correlations, or costs differs across states by enough that a decision taken under one would be wrong under the other. Remove any of the three and the word loses its content: without persistence there is only a mixture; without a difference in distributions there is nothing to detect; without hiddenness there is nothing to infer.[1]

The two-state hidden Markov model is the minimal formalization. In the stylized version used throughout this piece, the state is either calm or stressed. Each trading day the calm state gives way to the stressed state with probability 0.02 and the stressed state to the calm state with probability 0.08, so that calm spells last about fifty trading days on average, stressed spells about twelve, and the market spends roughly one day in five under stress. The daily return is drawn from a normal distribution with zero mean in both states; only the standard deviation differs, 0.8% in the calm state and 2.0% in the stressed one. Every parameter is chosen for exposition rather than estimated from anything.

Restricting the difference between states to the variance is deliberate. The regimes that matter most in practice are volatility regimes, because volatility is what risk models, position limits, and margin requirements all key off, and because it is the one moment of the return distribution that can be estimated with any precision over a horizon of days. A difference in means between regimes is very likely there and very nearly impossible to see at this horizon; a detector that leans on it will be led by noise. The variance is what can be seen, so the variance is what the model watches.

Filtering Versus Hindsight

Once the model is written down there are two different questions one can ask of it. The first: given every return observed up to and including today, what is the probability that today is a stressed day? That is the filtered probability, and it is the only object available to anyone deciding in real time. The second: given the entire history, including everything that happened afterward, what is the probability that a given past day was stressed? That is the smoothed probability, and it is what almost every published chart of market regimes shows, whether or not the caption says so.[2]

Figure 1 draws both for one simulated path of three hundred trading days. The shaded bands mark the true state, which the model never sees. The solid line is the filtered probability, updated one day at a time; the dashed line is the smoothed probability, computed after the fact from the full sample. Both are given the true parameters of the model, a favor no real detector receives.

Figure 1:  Real-Time and Hindsight Probabilities of the Stressed State on One Simulated Path300 trading days; shaded bands mark the true, unobserved stressed state
Stressed0.000.250.500.751.00050100150200250300Decision thresholdProbability of stressed stateTrading day
Filtered probability (uses returns through that day)Smoothed probability (uses the whole sample)

Note: Two-state hidden Markov model: the unobserved state moves from calm to stressed with probability 0.02 per trading day and back with probability 0.08 (expected spells of about 50 and 12 days). Daily returns are normal with zero mean and standard deviation 0.8% in the calm state and 2.0% in the stressed state. The filtered line is P(stressed | returns through day t) from the forward recursion; the smoothed line is P(stressed | all 300 days) from the forward-backward recursion. Both use the true parameters. Seeded pseudo-random draws.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

Three features recur in every path we have drawn. First, the filter is right most of the time and often a step behind. A stressed spell begins with a large return, and a single return of typical stressed size is not enough on its own: a 2% move occurs in the calm state a few times a year, so the filter, starting each day from a prior weighted heavily toward calm, commits at once only when the opening move is unusually large and otherwise needs another day or two. Here two of the three spells it catches are flagged on their first day, the long one near day 63 on its second. The exit is late for the mirror reason: a few quiet days inside a stressed spell look like the quiet days any stressed spell contains, so the filter lingers above the threshold for two or three days after each spell ends. Second, the filter raises false alarms. A single large calm-day return lifts the probability sharply, and it decays as ordinary returns arrive. Here two such alarms are raised near days 292 and 299, with no stressed spell behind them; each fades within a day or two, and on the day it was raised no real-time observer could have told it from a genuine onset. Third, the four-day spell near day 256 is never detected at all. Its probability peaks below 0.3; the spell ended before the evidence for it had accumulated.

The smoothed line has none of these defects, and that is the point. It hugs the bands, enters on time, and exits on time, because it is allowed to use the future. Whenever a chart of regimes looks clean, the reader should ask which of the two lines is on the page. At every point along the filtered line, the reader of Figure 1 knows more than the line does, and that asymmetry is the whole difference between explaining a regime and recognizing one.

How Late Is Late?

The lag in Figure 1 can be measured. Define the detection lag of a stressed spell as the trading days from its true onset to the first day the filtered probability reaches one half, and call the spell missed if it ends before that happens. Figure 2 shows the distribution of that lag over several hundred simulated spells, once for the model above and once for a less distinct version in which the stressed-state standard deviation is 1.4% rather than 2.0%.

Figure 2:  Distribution of Detection Lag for Stressed Spells, Two Stylized ModelsTrading days from true onset until the filtered probability first reaches 0.5
0%10%20%30%40%01234–56–10> 10MissedShare of stressed spells
Distinct states (0.8% vs 2.0%)Less distinct states (0.8% vs 1.4%)

Note: Lag is the number of trading days from the true onset of a stressed spell to the first day the filtered probability reaches 0.5; "Missed" counts spells that ended before it did. Computed over the 387 and 383 spells that begin while the filter is below threshold in two 25,000-day simulations of the model in Figure 1, the second with the stressed-state standard deviation lowered from 2.0% to 1.4% (the filter is given the matching parameters in each case).

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

In the stylized model, about two-thirds of spells are flagged within three trading days, and the median lag is a single day. The tail is long, and roughly one spell in eight is never flagged at all; nearly all of those lasted five days or fewer. That last group is not a failure of the detector. A regime that ends before its evidence arrives is invisible in principle, to any method, and a detector tuned to catch it would fire constantly. Public episodes illustrate the range of time scales without any need for numbers: an episode that resolves within minutes, as the flash crash of 2010 did, is beyond the horizon of any daily filter; one that builds over weeks, as the dislocation of March 2020 did, is within it.

The less distinct model makes the same point from the other side. Bring the two states closer together and the whole distribution shifts to the right: the median lag roughly triples, the tail thickens, and the missed share more than doubles. Nothing about the algorithm changed. The lag is a property of how well the observations separate the states, not of the recursion, and the way to shorten it is to find observations that separate them better: cross-sectional dispersion, option-implied volatility, order-book depth, and other series that respond to stress faster than a daily return does. The alternative is to act on a lower probability, and that has a price of its own.

False Alarms and Missed Regimes

The threshold of one half in Figures 1 and 2 was a convenience. It is in fact a policy choice, and the most consequential one in the design of any regime detector. Lower it and stressed spells are recognized sooner, at the cost of more calm days flagged. Raise it and the false alarms thin out, at the cost of slower recognition and more spells missed entirely. Figure 3 traces the trade-off, in the manner of a receiver operating characteristic, for the same two models: the share of stressed days correctly flagged against the share of calm days wrongly flagged, as the threshold sweeps from near zero to near one.

Figure 3:  The Trade-Off Between False Alarms and Missed Stress as the Threshold MovesShare of stressed days flagged against share of calm days flagged, thresholds from 0.02 to 0.98
0.000.250.500.751.000.000.250.500.751.00Threshold 0.2Threshold 0.5Threshold 0.8Share of stressed days flaggedShare of calm days flagged (false alarms)
Distinct states (0.8% vs 2.0%)Less distinct states (0.8% vs 1.4%)No information

Note: Each curve traces, for thresholds from 0.02 to 0.98 in steps of 0.02, the share of stressed days on which the filtered probability is at or above the threshold (vertical) against the share of calm days on which it is (horizontal), over the same simulations as Figure 2. The dashed diagonal is a detector with no information. Labeled points mark thresholds of 0.2, 0.5 and 0.8 on the distinct-state curve.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

No point on either curve is correct. The right threshold depends on the asymmetry between two costs: the cost of being late into a stressed regime, paid in drawdown, and the cost of responding to a false alarm, paid in turnover, in exposure given up during a calm period, and in the credibility of the next alarm. When the first dominates, the threshold should be low and the false alarms accepted as the premium on an insurance policy. When the second dominates, the threshold should be high and some lateness accepted. Either way the choice has to be made explicitly, with both costs in view, rather than left at whatever value the software defaulted to.

The second curve carries a harder lesson. When the states are less distinct, every point on the curve is worse at once: for any given false-alarm rate, fewer stressed days are caught, and no threshold rescues it. A regime that is real but not distinct enough to be detected in time is, for every practical purpose, not a regime. This is the argument for spending research effort on observations that separate the states, which moves the whole curve, rather than on threshold tuning, which only moves along it.[3]

What Counts as a Regime

Volatility is only one dimension along which markets change state. Figure 4 sets out four that matter for a systematic investor: what the calm and stressed states look like along each, which observations reveal the change, roughly how long it takes to see, and the characteristic way a detector built on that dimension is misled.

Figure 4:  Four Dimensions of Regime and How Each Reveals Itself
DimensionCalm stateStressed stateWhat reveals the changeTime to seeWhere a detector is misled
VolatilitySmall daily moves; thin tailsLarge moves that cluster; fat tailsReturns; option-implied volatilityDaysOne large calm-day move looks, for a day, like the onset of stress
CorrelationAssets driven by their own news; diversification worksOne factor dominates; pairwise correlations rise toward oneCross-sectional dispersion; short-window realized correlationDays to weeksRising correlation is usually a consequence of stress already under way, not a warning of it
LiquidityDeep books, narrow spreads; orders absorbedThin books, wide spreads; impact per unit traded risesOrder-book depth, quoted spreads, own trade impactHours to daysCan vanish and return within a session, too fast for a daily filter; looks normal until used
PolicyStable reaction function; announcements confirm expectationsReaction function shifts; announcements surpriseStatements, projected rate paths, tone of communicationOnset dated exactly; effect on returns takes monthsThe date is known at once; whether the return distribution changed is not known for months

Note: Qualitative summary of the mechanisms described in the text. “Time to see” is the order of magnitude of the delay between a change in the state and the point at which a real-time observer, using observations of the kind listed, can be reasonably confident of it.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

The dimensions differ most in their observability. Volatility can be read from returns alone within days. Correlation needs a cross-section, which is available daily, so it can be seen over short windows, but it tends to confirm stress rather than warn of it. Liquidity is visible intraday, faster than any of the others, and is also the most fleeting. Policy is the odd case: its onset is dated exactly, to an announcement, and its effect on the distribution of returns is the slowest of the four to establish.

In a crisis the four dimensions change state together, which is why crises are easy to name afterward. Outside a crisis they diverge: volatility can rise while liquidity is fine, and correlations can jump on a policy surprise with no change in volatility. A detector that waits for all four to agree will be confident only when confidence is no longer useful; one that combines them, weighting each by how fast and how reliably it speaks, will be earlier, at the price of more partial alarms.

Recognizing Rather Than Explaining

Five practices follow from taking the filtered line seriously. The first is to keep probabilities as probabilities: to size a response to the current probability of stress rather than to a binary label, so that a false alarm that decays over three days costs three days of partial caution rather than a full reversal. The second is to budget for false alarms in advance, deciding what rate is acceptable and holding the threshold to it, rather than discovering the rate in live trading. The third is to evaluate every detector on its filtered output, in event time, against episodes it did not see when it was built; a detector that looks good on smoothed output has been graded against the answer key. The fourth is to invest in observations that separate the states, the only way to move the curve in Figure 3 rather than slide along it. The fifth is to be candid about the spells no method can catch, and to hold enough general robustness that being late to a short one is survivable.

None of this makes the future legible. A regime explained afterward is a story; a regime recognized in time is a decision, and the distance between the two is measured in days. Those days are what the work is about.


  1. [1]The formalization follows Hamilton (1989), “A New Approach to the Economic Analysis of Nonstationary Time Series and the Business Cycle,” Econometrica, which introduced the Markov-switching model in the context of output growth. Its application to the variance of asset returns rather than to the mean is the version most useful at the horizon considered here.
  2. [2]The filtered probability is P(S_t = stressed | r_1, …, r_t) and the smoothed probability is P(S_t = stressed | r_1, …, r_T) for a sample ending at T. The two are computed by the forward and forward-backward recursions respectively. In-sample fits of regime models, and the charts drawn from them, report the second.
  3. [3]Ang and Bekaert (2002), “International Asset Allocation with Regime Shifts,” Review of Financial Studies, document that international equity returns are well described by a regime in which volatility and correlation rise together, and examine what an allocator who takes such regimes into account, with the lag that real-time inference implies, can and cannot gain from doing so.

Interested in related insights?

Correlation Is Not Constant: What Happens to Diversification When the Regime Changes?

What Volatility Knows: The Market's Forecast of Its Own Uncertainty, and How to Read It

Enjoyed this piece?

Share your thoughts!

This document is provided for informational purposes only and does not constitute investment advice or an offer to sell (or the solicitation of an offer to buy) any security, investment product, or service.

The views expressed are those of OAK ST LLC as of the date of the document, are subject to change without notice, and may not reflect the criteria used by OAK ST LLC to evaluate investments. Figures described as illustrative, stylized, or simulated are hypothetical constructions prepared for exposition; they do not depict the results of any OAK ST LLC strategy, portfolio, or account, and no representation is made that any account will or is likely to achieve results similar to those shown. Historical market trends are not reliable indicators of future market behavior.

Information obtained from third-party sources is believed to be reliable but has not been independently verified, and OAK ST LLC does not guarantee its accuracy or completeness. Nothing in this document is a recommendation to buy, sell, or hold any instrument.

This document may not be reproduced or distributed without the prior written authorization of OAK ST LLC. The Terms of Use and the Important Legal and Regulatory Disclosures govern its use. Copyright © 2026 OAK ST LLC. All rights reserved.