Autonomy: How Far an Investment System Can Run Itself, and Where the Lines Should Not Move

An investment system does not become autonomous all at once, and it should not. The vocabulary that has proved useful for driving automation makes the point. A car with adaptive cruise control and a car that drives itself through a city are separated not by one capability but by a ladder of them, and the rung that matters most is not the one where the machine takes the wheel. It is the one where the machine becomes responsible for noticing that it should hand the wheel back. The same is true of a quantitative investment process. A model can produce a forecast, an optimizer can turn forecasts into a portfolio, and an algorithm can work an order without a person touching it, and none of that settles the question of whether the people can step away. What settles it is whether the system knows when it is outside the conditions it was built for, and what it does then.

This piece sets out a framework for thinking about that question. It borrows the levels-of-autonomy structure from driving and applies it separately to each stage of a closed loop that runs from research through validation, portfolio construction, risk control, execution, and monitoring to self-diagnosis and back to research. Three arguments follow. Readiness differs across the stages of the loop by more than most discussions of automation in finance acknowledge. The practical limit on autonomy is not the sophistication of the models but the arithmetic of review: how many decisions a person can examine with care, and how many cycles a fault can run before anyone looks. And a small set of boundaries should stay under human control regardless of level, not because the machine is distrusted but because a bounded failure surface is what makes delegation reasonable in the first place. Throughout, we describe how Oak St. thinks about the problem, not what any system of ours has done.

Levels, Not a Switch

The taxonomy used for driving automation distinguishes six levels by two questions: who performs the task, and who is responsible for monitoring the environment and deciding when the task can no longer be performed safely.[1] At the bottom, a person does everything and the machine warns. In the middle, the machine steers and brakes within a narrow envelope while a person watches every moment. Above that, the machine also does the watching, within conditions it can describe, and hands control back when those conditions end. At the top there are no conditions. The structure transfers to investing with little strain, because each stage of an investment process is also a task with an operating envelope, a set of conditions under which its assumptions hold, and a question of who notices when they stop holding.

Figure 1 sets out the six levels as we apply them. The definitions are deliberately about responsibility rather than technology. A stage is at Level 2 not because its models are simple but because a person is expected to review each action before or immediately after it happens; it is at Level 3 not because its models are sophisticated but because the system itself is expected to detect that it has left its envelope, and a person is expected to be available when it does. Level 5, unrestricted operation, is included for completeness. We do not regard it as a destination for an investment process, because the operating domain of a market is not closed: the conditions under which any model was built can and will end, and a system that claims to need no envelope is a system that has not described its own.

Figure 1:  Six Levels of Autonomy, Applied to a Stage of an Investment ProcessDefined by who acts and who is responsible for noticing; the driving analogy is a mnemonic, not a claim of equivalence
LevelNameWho actsWho noticesDriving analogy
Level 0ManualA person; tools computeThe personNo automation; warnings at most
Level 1AssistedA person, with suggestions from the systemThe personLane-keeping or braking assist
Level 2SupervisedThe system, within a narrow envelopeThe person, on every actionSteering and speed automated; hands on the wheel
Level 3ConditionalThe system, within an envelope it can describeThe system; a person on call to take overEyes off the road; ready to resume on request
Level 4DomainThe system, including its own degraded modesThe system; people define the domainDriverless within a mapped area
Level 5UnrestrictedThe system, anywhereThe systemDriverless anywhere; not a destination for an investment process

Note: A qualitative framework adapted from the levels of driving automation. Each level is a property of one stage of the process, not of the process as a whole; the same firm can, and in our view should, run different stages at different levels.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

Two features of the ladder are easy to miss. First, the step from Level 2 to Level 3 is the largest, because it is the step at which the burden of noticing moves from the person to the machine. Everything below it can be built with models that are wrong in unknown ways, because a person is checking; everything above it requires models that know something about how they are wrong. Second, the levels are properties of stages, not of firms. A process can, and in our view should, be at different levels in different places, and the interesting design work is in deciding where the levels differ and why.

The Loop and Its Checkpoints

A quantitative investment process is best drawn as a loop rather than a pipeline, because its last stage feeds its first. Research proposes hypotheses. Validation decides which of them survive an honest out-of-sample test, adjusted for the number of trials it took to find them.[2] Portfolio construction turns surviving forecasts, a risk model, and a cost model into target weights. Pre-trade limits check those weights against bounds that were set in advance. Execution turns the difference between the current and target portfolio into orders and works them against available liquidity. Monitoring compares what happened with what was expected: fills, slippage, exposures, the behavior of the forecasts themselves. Self-diagnosis asks whether the machinery is healthy: whether data arrived, whether it looked like the data the models were trained on, whether latencies and error rates are in their usual ranges, whether a model's live behavior has drifted from its validated behavior. And review carries what the live evidence says back into research, where it either strengthens a hypothesis or retires it.

Figure 2 draws the loop and marks, at each stage, who is responsible for the action and who is responsible for noticing. The pattern is the point. Human sign-off sits at two places: the promotion of a hypothesis into something that trades, and the reading of live evidence back into the research agenda. Human availability, a person on call and reachable, sits at monitoring. The machine acts alone in the middle of the loop, but only inside limits that it did not set, and self-diagnosis is given exactly one authority of its own: it may pause or reduce, never extend.

Figure 2:  The Closed Loop, with the Human Checkpoints MarkedEight stages; the last feeds the first. The amber label records who acts and who notices at each stage
ResearchHypotheses proposed, testedMachine proposesValidationOut-of-sample, trial-adjustedHuman signsPortfolio constructionForecasts and risk to weightsMachine, within limitsPre-trade limitsBounds set by people, in codeMachine enforcesExecutionOrders paced to liquidityMachine, within limitsMonitoringFills, slippage, exposuresHuman on callSelf-diagnosisData, latency, drift checksMachine; may only pauseReviewLive evidence to researchHuman signs

Note: The loop is drawn in the order in which a decision propagates; review feeds back into research. “Machine, within limits” means the stage acts without review against bounds it cannot change; “human signs” means a person must approve before the loop proceeds; “human on call” means a person must be reachable and able to intervene. The assignment is a design position, not a description of any particular system.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

Why do the human checkpoints cluster at the boundaries between stages rather than inside them? Because a stage is a well-posed problem with an envelope that can be written down, and the boundary is where the assumptions of one stage become the inputs of the next. An optimizer that receives a forecast does not know whether the forecast was validated honestly; an execution algorithm that receives an order does not know whether the target weights made sense. Errors that cross a boundary compound, because each subsequent stage treats them as facts. A person placed at a boundary is placed where the compounding can be interrupted. A person placed inside a stage is usually placed where the machine is already better at the task than the person, and the checkpoint becomes a ritual.[3]

Readiness Is Uneven Across the Loop

The stages of the loop are not equally ready to climb. Execution has been substantially automated across the industry for two decades, and the reasons are instructive: its objective can be fully specified in advance, its feedback is fast and unambiguous, its failures are bounded by the size of the order, and its operating envelope, a description of the liquidity conditions under which a schedule makes sense, can be written down and checked. Risk control is automated in enforcement and human in specification, which is the right division and is not likely to change. Portfolio construction sits in the middle: the optimizer acts, but the objective it optimizes embodies judgments about risk aversion, cost, and constraint that people revise. Research is the least automated stage in the sense that matters. A system can generate and test candidate hypotheses at a scale no group of people can match, but deciding what a surviving hypothesis means, whether there is an economic reason it should hold, and whether it is a new idea or an old one wearing different data, remains human work.

Self-diagnosis is the surprise. It is the stage most people would rank as an engineering detail, and it is the one whose readiness bounds every other. Figure 3 makes the unevenness visible with a stylized readiness score for each stage at each level. The scores come from a simple model, stated in the note, in which each stage has a level at which its readiness is one half and readiness falls smoothly above it; the centers are our qualitative reading of current practice across the industry, not a measurement of any system. The ordering, not the decimals, is the claim.

Figure 3:  Stylized Readiness to Operate at Each Level, by Stage of the LoopA logistic readiness curve per stage; illustrative
Level 0Level 1Level 2Level 3Level 4Level 5Research94.8%74.9%32.6%7.27%1.26%0.206%Validation98.2%89.9%59%18.9%3.65%0.611%Portfolio construction98.9%93.9%71.3%28.7%6.14%1.05%Risk control99.7%98.2%89.9%59%18.9%3.65%Execution99.9%99.1%94.8%74.9%32.6%7.27%Monitoring99.5%96.9%83.7%45.5%11.9%2.15%Self-diagnosis92.7%67.4%25.1%5.17%0.877%0.143%

Note: Readiness is 1 / (1 + exp((L − c) / w)) with w = 0.55 and a stage-specific center c at which readiness is one half: research 1.6, validation 2.2, portfolio construction 2.5, risk control 3.2, execution 3.6, monitoring 2.9, self-diagnosis 1.4. The centers are a qualitative assessment made for exposition, not a measurement of any system.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

What makes a stage ready is a short list. Its objective must be specifiable before the fact, so that the machine is not being asked to infer what it should want. Its feedback must be fast and unambiguous, so that errors are visible within the horizon at which they can still be corrected. Its failures must be bounded and, ideally, reversible. And its envelope must be describable, so that the machine can be told, in terms it can check, when it is outside the conditions it was built for. Execution meets all four conditions. Research meets none of them well: its objective is contested, its feedback arrives over years, its failures are unbounded in the sense that a bad idea can consume any amount of capital, and its envelope is the entire future. That is not an argument against automating parts of research. It is an argument that the level at which research is automated should be set by these four conditions and not by what the tooling can do.

The Arithmetic of Delegation

Discussions of autonomy in finance tend to be about capability: whether a model is good enough to act on its own. We think the binding constraint is somewhere else. It is the arithmetic of review. A person can examine a bounded number of decisions in a day with genuine attention, and that number does not scale with the number of decisions a system makes. As a process climbs the ladder, decisions multiply: more names, more frequent forecast updates, more order slices per rebalance. At some point the share of decisions a reviewer can examine falls below the share needed for review to mean anything, and human oversight becomes sampling. Sampling is not worthless, but it is a different activity from supervision, and a firm that has not noticed the difference is running at a higher level than it believes.

Figure 4 puts illustrative numbers on that transition for a stylized Level 3 process. The values are hypothetical and chosen to make the arithmetic legible; the conclusion does not depend on them. Once a reviewer can see a fraction of a percent of what the system decides, the safety of the process can no longer come from the reviewer. It has to come from the limits every decision passes through and from the system's ability to detect its own faults.

Figure 4:  When Review Becomes SamplingA stylized Level 3 process; hypothetical values
  • 20,000

    Decisions per day in the stylized process

    200 names, 10 forecast updates, 10 order slices each

  • 60

    Decisions a careful reviewer can examine per day

    eight hours at eight minutes each

  • 0.3%

    Share of decisions that can be reviewed

    review has become sampling

  • 100%

    Share that pass through hard limits

    every decision, in code, before it acts

Note: Decisions per day is 200 × 10 × 10; reviewer capacity is 480 minutes at 8 minutes per decision; the share reviewed is their ratio. All values are hypothetical and chosen to make the arithmetic legible, not measured from any process.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

The second half of the arithmetic concerns time rather than volume. At low levels, a fault is caught at the next human look, which is the next cycle. As the level rises, the number of cycles between human looks grows, and it grows geometrically, because each rung of the ladder removes a category of routine inspection rather than a fixed amount of it. A fault that would have run for one cycle at Level 0 can run for hundreds at Level 4 before anyone is scheduled to notice. Figure 5 draws that growth in a stylized model in which the cycles between human looks quadruple with each level, and then shows what automated self-diagnosis does to it. With even a modest per-cycle probability of detecting a fault, the expected number of cycles a fault survives stops growing with the level and settles at a ceiling set by the detector, not by the ladder.[4]

Figure 5:  Expected Cycles a Fault Survives, With and Without Self-DiagnosisCycles between human looks quadruple with each level; log scale
1101001000Level 0Level 1Level 2Level 3Level 4Level 5Expected cycles before detectionAutonomy level of the acting stage
Human looks only (p = 0)Self-diagnosis, p = 0.02 per cycleSelf-diagnosis, p = 0.2 per cycle

Note: Cycles between human looks are n = 4^L at level L. With a per-cycle detection probability p, the expected cycles a fault survives is (1 − (1 − p)^n) / p; with p = 0 it is n itself. The quadrupling and the two values of p are chosen to illustrate the mechanism, not estimated from any system.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

This is the sense in which self-diagnosis bounds everything else. The level at which a process can safely act at any stage is limited by the level at which it can diagnose itself, because the alternative is a growing interval during which a fault runs unobserved. A system that acts at Level 3 and diagnoses itself at Level 1 is not a Level 3 system; it is a Level 1 system with fast hands and a long interval between looks. This is also why the readiness ordering in Figure 3 should be read from the bottom up: the stage least ready to climb is the one that determines how far the others may.

Boundaries That Do Not Move

None of the above is an argument for distrust. It is an argument for a bounded failure surface, and the instrument for that is a set of boundaries that do not depend on the level of autonomy at all. A boundary in this sense is a constraint that is enforced in code on every decision, that the system cannot relax, and that can be changed only by people, through a process with a record and a delay. Figure 6 lists the boundaries we regard as belonging to that set. The column that matters most is the last one, which answers the same way in every row.

Figure 6:  Boundaries That Do Not Move, Whatever the LevelConstraints enforced on every decision, changed only by people, never relaxed by the system
BoundaryWhat it boundsEnforced byChanged byRelaxable by the system?
Exposure limitsGross, net, per-name, and per-factor exposurePre-trade check, in codeTwo people, with a written reason and a waiting periodNo
Loss limitsThe loss, daily and cumulative, at which trading haltsAutomatic haltRisk oversight, on the recordNo
Order throttlesOrder rate, size, and price collars, per venueGateway, before the wireTwo people, with a written reason and a waiting periodNo
Instrument and venue listsWhat may be traded, and whereWhitelist, in codeCompliance sign-offNo
Model promotionWhich models may direct live capitalDeployment gateHuman sign-off after validationNo
Data provenanceWhich inputs may reach a modelLineage check at ingestionThe data's owner, on the recordNo
Limit configurationThe limits themselvesVersioned, signed configurationPeople only, with a waiting periodNo
Kill switchEverythingAn independent path to a flat bookHuman onlyNo

Note: A design position on the set of constraints that should sit outside the autonomy ladder entirely. “Relaxable by the system?” asks whether any automated component may loosen the boundary on its own authority; tightening, pausing, and refusing remain available to the system at every level.

Sources: Oak St. research. Illustrative, stylized simulation prepared for exposition; not derived from any Oak St. portfolio, strategy, or live data.

Three design rules follow from the table. The first is asymmetry. The system may tighten any boundary on its own: it may reduce exposure, slow its order rate, or stop, and it should be built so that doing so is easy and cheap. It may never loosen one. A machine that can only move in the safe direction on its own authority is a machine whose worst autonomous act is to do too little, which is a failure a firm can afford. The second is independence. The path from a person to a flat book must not run through the system being stopped: a kill switch that depends on the health of the thing it kills is not a kill switch. The third is legibility. Every autonomous action must leave a record that a person can read afterward and that says what the system saw, what it decided, and which limits it checked. Autonomy without that record is not delegation; it is abdication, and it is the version of autonomy that regulators, investors, and the people who run the firm are right to be wary of.

Governance is the process around the boundaries, and it is worth being concrete about its shape. A change to a boundary should require two people, a written reason, and a waiting period long enough that the change cannot be made in the middle of the event that motivated it. A change to the level at which a stage operates should be a decision with a record, made against the four readiness conditions, and reversible on the same day it is made. And the people at the checkpoints in Figure 2 must have the time and the information to refuse. A sign-off that cannot be withheld is not a checkpoint, and a person on call who cannot understand what the system is doing is not oversight. The most common way for a process to end up at a higher level than it admits is for its checkpoints to remain on the diagram after they have stopped being real.

How We Think About It

Our own view of autonomy follows from the framework rather than the other way round. We treat the levels as a description of where each stage of a process is, not as an ambition for where it should be; the aim is a system whose level at every stage matches what it can demonstrate, and a stage is promoted one rung at a time, with evidence, and with a record of who decided. We invest in self-diagnosis ahead of acting capability, because the interval between looks is the quantity that governs how much acting capability is safe to use. We build the asymmetry in from the start: the system may reduce, pause, or refuse on its own; it may propose a change to a limit; it may not make one. And we keep the human checkpoints at the stage boundaries and try to design them so that refusing is possible, which means giving the people at them fewer decisions to examine and better information about each.

The word autonomy suggests independence. In an investment system it should mean something narrower and more useful: a well-described envelope, a machine that knows when it is inside it, and people who decide where the envelope's edges are and who can move them only slowly. That is less dramatic than a system that runs itself, and it is the version worth building.


  1. [1]The structure is adapted from the levels of driving automation in SAE International's J3016 standard, which distinguishes the levels by who performs the driving task and who is responsible for monitoring the driving environment and responding when the automated system cannot. We borrow the shape of the taxonomy, not its specifics, and make no claim that the analogy is exact.
  2. [2]An automated research loop that generates and tests many candidate hypotheses is, unless the validation stage corrects for the number of trials, a machine for producing false discoveries. Harvey, Liu, and Zhu (2016) and Bailey and López de Prado (2014) set out how large the correction must be; the sign-off in Figure 2 is where a person confirms that it was applied.
  3. [3]Bainbridge (1983) observed that automating the routine parts of a task leaves the person with the parts that are hardest and least practiced, and asks them to intervene precisely when the automation has failed in a way it could not describe. Placing checkpoints at stage boundaries, where the question is whether to accept an input rather than how to perform a task, is partly an answer to that observation.
  4. [4]If a fault is detected independently on each cycle with probability p, the number of cycles until detection is geometric, and its expectation truncated at n cycles, the interval between human looks, is (1 − (1 − p)^n) / p. That quantity is what Figure 5 plots; for p = 0 it reduces to n itself.

Interested in related insights?

One Portfolio, Thousands of Decisions: Why Turning Independent Forecasts into a Single Set of Positions Is a Hierarchy, Not a Sum

Prediction Without Explanation: How Much to Trust a Model That Is Right for Reasons No One Can State

Enjoyed this piece?

Share your thoughts!

This document is provided for informational purposes only and does not constitute investment advice or an offer to sell (or the solicitation of an offer to buy) any security, investment product, or service.

The views expressed are those of OAK ST LLC as of the date of the document, are subject to change without notice, and may not reflect the criteria used by OAK ST LLC to evaluate investments. Figures described as illustrative, stylized, or simulated are hypothetical constructions prepared for exposition; they do not depict the results of any OAK ST LLC strategy, portfolio, or account, and no representation is made that any account will or is likely to achieve results similar to those shown. Historical market trends are not reliable indicators of future market behavior.

Information obtained from third-party sources is believed to be reliable but has not been independently verified, and OAK ST LLC does not guarantee its accuracy or completeness. Nothing in this document is a recommendation to buy, sell, or hold any instrument.

This document may not be reproduced or distributed without the prior written authorization of OAK ST LLC. The Terms of Use and the Important Legal and Regulatory Disclosures govern its use. Copyright © 2026 OAK ST LLC. All rights reserved.