Forecast reliability rates how much to trust a forecast's confidence band — Calibrated, Wider range, Uncertain or Pending — from that ticker's own backtested band calibration.
Calibrated — the 90% band has historically covered close to 90% of realized prices for this ticker, so you can read it at face value. Wider range — the band is softer than our best-calibrated forecasts; lean toward the wide ends of the range, not the point estimate, and size smaller.
Uncertain — a real, backtested reason to distrust the band (historically over-confident, or unusually wide); treat the call as WAIT until it clears. Pending — the stock is too newly listed to backtest the band yet, so it is "not enough evidence," not a red flag.
Reliability is derived only from numbers the model already computes: the horizon's backtested CI90 coverage (does the 90% band historically contain about 90% of outcomes?), how wide the band is relative to price, and whether a calibration haircut had to widen the raw band. It is a statement about this ticker's own historical band calibration — the per-ticker companion to the platform-wide CI90 coverage figure — never a promise about any single future trade.
Reliability is per-ticker because any headline figure hides wide dispersion. On the committed TOP-20 backtest the 90% bands covered 78–95% of realized prices at 3 months depending only on which quarter the forecast started in — and individual windows are worse still: UNH started 30 Jun 2025 covered 10% at 3 months, and NVDA started 31 Dec 2024 covered 2% at 1 year (badly over-confident — exactly what the label flags). A quiet, familiar name is no guarantee either. Check the label on the name you trade, not any platform average.
We publish no coverage percentage for an individual ticker, because we have not measured one. A single ticker's backtest covers one hold-out window — the model fitted up to one past date, scored against the one stretch of prices that followed — so the number swings on nothing but which date it started from, and it is computed without the adjustments the forecast you actually read applies. A ticker page shows which WAY the range missed instead of a figure, because both ends are misses: a 90% range that held almost every price is too wide to be useful, not a good score. The measured record is the universe table on the calibration page, reported separately at each quarterly forecast start date. See the full AAPL forecast for the current band this reliability label attaches to.
A well-calibrated band means the width of the uncertainty is honest — not that the forecast pointed the right way. How often the model calls up-vs-down is a separate axis (directional hit rate), and how strong its current view is (model confidence) is a third, distinct number. Reliability is only about whether the band is honest about its own uncertainty.
No. Reliability only rates how much to trust the band's width — it is separate from the forecast's direction and its entry timing. A calibrated band can still point the wrong way.
Treat the call as WAIT until it clears — Uncertain means there is a real, backtested reason to distrust the band, so acting on the point estimate is a bet you can't size honestly.
Because calibration varies a lot per ticker. On the committed TOP-20 backtest, AAPL's 3-month band covered 85.7% of realized prices while AVGO's 1-year band covered only 36.5%. A structural break is one cause, but not the only one — some large, stable names are poorly covered too, which is why the label is computed per ticker rather than assumed from the company.
AAPL analysis shows this metric in context, or browse all S&P 500 tickers.
Educational research only — not investment advice.