Short answer: our forecast bands are honestly calibrated as of the most recent recalibration, but calibration is not the same question as "does it predict direction correctly" — and we publish both numbers, including a past mistake we corrected in public.
An earlier headline claimed 97% CI90 accuracy on our TOP-20 backtest. That figure was a single curated snapshot that was not reproducible from the committed backtest artifact — we retracted it. Honest re-testing at that time showed real overconfidence: roughly 74.7% (3-month), 93.3% (6-month), and 80.8% (1-year) coverage against a 90% target — two of three horizons meaningfully below where a well-calibrated model should sit. Those are historical readings, not the current ones. We shipped a horizon-dependent band recalibration rather than quietly loosening the bands, and published the before/after difference. On the current committed backtest, measured against the model as served across nine quarterly forecast start dates, coverage ran 78–95% of realized prices at 3 months, 71–92% at 6 months and 86–96% at 1 year — published per start date, never as one average.
These figures describe the percentage of realized prices that landed inside the published 90% range, not whether the model's directional call was correct. We never present a forecast mean as a standalone buy/sell signal, and per-ticker dispersion around the aggregate is still wide — individual names range from the mid-30s to nearly 100% coverage, and that spread did not narrow when the aggregate improved.
See the live calibration page for the current per-horizon figures, and the live track record for closed-trade outcomes, both self-updating and both showing the losses alongside the wins.
We hold ourselves to a pre-registered well-calibrated range of 85–95% coverage for a 90% band, and report honestly when a horizon sits outside it. A horizon can miss in either direction: "over-confident" means the band is too narrow to contain what actually happened, "under-confident" means it is wider than needed and carries little information. Either way a recalibration is tracked openly. The current per-horizon figures are not restated here, because they update as the backtest window rolls forward — read them, with the full per-ticker breakdown and the misses included, on the live calibration page.
A forecasting product that only shows favorable numbers is not trustworthy, and trust is the entire basis on which anyone would pay for a forecast. The live calibration page and the live track record update automatically as the backtest window rolls forward and as trades close — both show the losses alongside the wins, not a cherry-picked highlight reel.
Our 90% confidence bands are honestly reported. On the current committed backtest, measured against the model as served across nine quarterly forecast start dates from 31 Mar 2024 to 31 Mar 2026, coverage ran 78–95% of realized prices at 3 months, 71–92% at 6 months and 86–96% at 1 year. We publish every start date rather than one average, and individual windows are worse still. That is a narrower claim than "the direction call is usually right".
Calibration measures whether a stated 90% band actually contains 90% of realized prices. Directional accuracy measures whether the forecast correctly called up vs. down — a model can be well-calibrated while directional hit-rate is only modestly better than a coin flip.
Yes — an earlier "97% CI90 accuracy" headline was a non-reproducible curated snapshot, which we retracted, corrected by a public band recalibration.
The calibration page and track record page both self-update from the committed backtest artifact and the closed-trade aggregate.
Browse all S&P 500 tickers to see this metric applied to individual companies.
Educational research only — not investment advice.