Where our model fails: structural breaks

Aggregate calibration figures can hide the names where a model does poorly, so this page names ours: AVGO at the 1-year horizon, whose band covered about 36.5% of realized outcomes on the committed backtest — by far the worst cell in the table. It illustrates the underlying failure mode — a structural break the model, built entirely on historical price paths, had no way to anticipate.

The aggregate number hides the worst cases

The headline coverage figures on the calibration page are averages across a basket of tickers. An average close to target can coexist with a handful of names that are badly miscalibrated, because good coverage on most tickers offsets poor coverage on a few. That is exactly the pattern at the 1-year horizon: the aggregate reads as well-calibrated while AVGO, GOOGL and MSFT, specifically, are not.

We publish the per-ticker breakdown for this reason — an aggregate number alone would let the worst cases hide in plain sight.

AVGO: a rally the historical path never suggested

AVGO went through a genuine re-rating — an AI-infrastructure demand surge and a major acquisition changed what the market was willing to pay for the business — that took its price well outside the range its own trailing history would have implied a year earlier. A model that forecasts from historical path statistics has no mechanism to anticipate a break of that kind: it can only widen or narrow a band around what the past distribution suggests is plausible, and a structural break is precisely a departure from that distribution. The result at the 1-year horizon is a band that covers a minority of the realized path.

Current figures are on the calibration page, filterable per ticker.

BRK-B: the same failure, the opposite direction — and a recovery

BRK-B showed the mirror case: the model overestimated where price would land, and the band's center sat above the eventual outcome for much of the horizon. The mechanism is the same underlying problem — historical path statistics assume the future resembles the recent past closely enough to bound it, and a name that breaks that assumption produces a band that's calibrated to the wrong market conditions.

On the current committed backtest, BRK-B has recovered: its 1-year band covered about 88.5% of realized outcomes, close to the 90% target and inside the well-calibrated range. We leave the post-mortem up rather than deleting it, because a failure that resolved is still part of the record.

Two different companies, two different directions of error, one shared cause: history-based path forecasting is blind to breaks from the pattern it was trained on. The weak 1-year names on the current record are AVGO at about 36.5%, GOOGL at 56.7% and MSFT at 75.0% — and a poorly covered name is not always one with an obvious structural break, which is why the per-ticker table, not a story, is the thing to check.

Honesty about the recalibration: wider bands, not smarter forecasts

A horizon-dependent band recalibration brought the aggregate coverage figures back toward their targets across all three horizons. It is important to say plainly what that fix actually was: the bands got wider, giving more room for realized prices to land inside them. Coverage improved because the acknowledged uncertainty grew, not because the underlying forecast got better at predicting where prices would actually go. That is a meaningful improvement in honesty — the stated confidence level now matches reality more closely — but it is not a claim that the model became smarter about AVGO or the next name to have a structural break.

What this means if you use the forecast

A wide band on a specific ticker is not decoration — treat it as a signal that history-based forecasting has less to say about that name, not as noise to look past. And do not read good aggregate calibration as a guarantee about any one ticker: check the per-ticker table before trusting a band on a name that has been through a recent change in its own market conditions.

Where this fails

This page is a post-mortem, not a fix. The model does not predict structural breaks — it cannot, by construction, since it forecasts from historical price-path statistics, and a structural break is precisely the kind of event that breaks the pattern those statistics describe. Recalibrating band widths after the fact improves coverage honesty; it does not give the model foresight it never had.

See the evidence

Frequently asked questions

Educational research only — not investment advice.