What does the "Model Quality" panel actually measure?

Model Quality is a per-ticker report card on our own forecast. Before you read what the model predicts for a stock, this panel tells you how well that same model did when it was made to forecast that same stock's recent past with the recent past hidden from it. A confident forecast on a ticker the model has historically handled badly is worth less than a hesitant one on a ticker it handles well.

Where the numbers come from

Every row in the panel comes from the same honest exercise, run fresh for that ticker during the scan. We cut the price history short at a point in the past, re-fit the model using only the data available up to that cut, let it forecast forward into a period it has never seen, and then compare its forecast against what the stock actually did. This is called a walk-forward test, and it is the only kind of test that means anything: a model scored on data it was fitted to will always look brilliant.

The panel shows one column per forecast horizon, because a model can be reliable over three months and useless over a year. Where a value could not be computed you will see a dash — we would rather show nothing than invent a number.

The six rows, one at a time

Growth is the model's central forecast at that horizon, in percent — the headline prediction itself, repeated here so you can read it next to the quality of the machinery that produced it. The chance of a gain is the model's estimated probability that the price finishes above where it is now. Neither of these is a quality measure; they are the forecast. The other four are the report card. Read the full treatment of the forecast itself in expected growth.

CI90 is the coverage of the 90% confidence band in the walk-forward test: of the days in the hidden period, what share saw the real price land inside the band the model drew? A well-behaved 90% band should contain the truth about 90 times out of 100. The panel colours this green at 90 and above and red below 70. See CI90 coverage for the full treatment.

MAPE measures how far the central forecast finished from the real price, as a percentage of that real price. Be precise about what our panel computes here: it is the error at the very end of the hidden window — |forecast − actual| ÷ actual × 100 — not an average of the errors along the whole path, which is what the name MAPE usually implies elsewhere. Green at 15 or below, red above 40. Background: MAPE.

CI width is the average width of that 90% band, expressed as a percentage of the price at the start of the test. This row exists to stop CI90 from lying to you. A model can score perfect coverage by predicting "somewhere between $10 and $500" — always right, always useless. Green at 50 or below, red above 150. Read CI90 and CI width together or do not read either.

Winkler is the one number that refuses to be gamed, because it combines both of the above: it rewards a narrow band and penalises the band for every time the real price fell outside it, with the penalty scaled by how far outside. A model cannot win by being vague and cannot win by being narrow and wrong. Lower is better; green at 50 or below, red above 150. Full treatment: the Winkler score.

How to actually read the panel (illustrative)

Example (illustrative — invented numbers chosen to show how the rows interact, not our results). Ticker A comes back with CI90 = 96 and CI width = 180. The coloured coverage looks excellent. It is not: a band that spans 180% of the starting price will contain almost anything. The Winkler row would be red, and Winkler is the row telling the truth.

Ticker B comes back with CI90 = 74 and CI width = 22. Coverage is amber-to-poor — the band missed more often than a 90% band should. But the band is tight, and if the misses were near misses, the Winkler score can still be respectable. This is a model being confidently approximately right, which is a very different failure from being vaguely right.

The reading habit worth building: never look at one row. Coverage tells you whether the uncertainty is honest, width tells you whether it is useful, MAPE tells you whether the centre was close, and Winkler arbitrates between them.

When Model Quality misleads you

It is a sample of one. This is a single truncated re-run on a single stock. A green CI90 here is not evidence the model is calibrated on that stock — it is one draw from a distribution. Our published calibration pools the whole universe precisely because a per-ticker number cannot carry that weight, and the t-statistic page explains why a handful of observations proves nothing.

It scores the recent past, and the recent past may not resemble the next few months. A stock that was quiet through the test window will make almost any model look calibrated. A stock that had an earnings shock or a structural break will make a good model look broken. The panel cannot tell those apart, and neither can the colour of the cell.

And it is a measure of the model, not of the stock. Nothing in this panel says a company is a good investment, and nothing in it says the forecast will be right this time. It says: here is how this machinery behaved last time it was pointed at this stock with the answer hidden. That is a reason to weight a forecast, never a reason to trust it. On why every model of this kind can look better than it is, read overfitting.

Frequently asked questions

What is the Model Quality panel?

A per-ticker report card on our own forecast. We hide the recent price history, re-fit the model on what remains, let it forecast into the hidden period, and score how it did — on band coverage, band width, central error and the Winkler interval score.

Is a high CI90 in this panel a good thing?

Only when the CI width row is also reasonable. Coverage alone is trivial to score well on: a band wide enough to contain any plausible price will have near-perfect coverage and tell you nothing. Read coverage and width together, and let the Winkler score arbitrate.

Why is MAPE here not a path average?

Because in this panel it is deliberately the error at the end of the hidden window — how far the central forecast finished from the real price, as a percentage of that price. The name is the conventional one; the computation is stated here so it cannot be misread.

Does a green Model Quality panel mean the forecast will be right?

No. It means the model handled this stock reasonably in one walk-forward test. That is a single observation, on a past that may not resemble the future. Use it to weight a forecast, never to trust one.

See it on a ticker

Browse all S&P 500 tickers to see this metric applied to individual companies.

Related terms

Educational research only — not investment advice.