Any tool can produce a beautiful backtest. The hard question is whether it tells you anything about the future or just describes the past really well — and a flawless track record is usually the warning sign, not the selling point.
In-sample means the tool was tested on the same historical data used to build and tune it. Out-of-sample means it was tested on data it never saw during tuning — a fair simulation of meeting the future for the first time. An in-sample result is like grading a student on the exact questions they were given the answers to; an out-of-sample result is a real exam. Only the second one predicts how the tool does on your money.
A model overfits when it’s flexible enough to fit every wiggle of the historical data — including the random noise that won’t repeat. It ends up memorising the past rather than learning a pattern that generalises. The tell is a stark gap: a spectacular in-sample backtest that falls apart out of sample. Example (illustrative): a strategy showing a flawless 100% win-rate on the exact five years it was built on, then barely breaking even on the next two years it never saw, is overfit. The perfect history was fitted, not forecast.
No out-of-sample or walk-forward test: if the tool only shows results on the data it was built with, treat the numbers as decoration, not evidence. Suspiciously perfect results: real markets are noisy, so a track record with no losing periods usually means the noise was fit, not the signal. Many knobs, little data: a model with dozens of tunable parameters and only a few years of history has enough freedom to fit almost any past. A conveniently chosen date range: results that only shine on one hand-picked window (and quietly skip others) are a red flag — the next lesson covers this cherry-picking in depth.
The gold standard is walk-forward validation: tune the model on an early stretch of history, test it on the next unseen stretch, roll the window forward, and repeat. It mimics how the tool would have actually been used over time, so the reported accuracy reflects genuine out-of-sample performance. When a tool reports backtest accuracy or an error metric like MAPE, the number is only meaningful if it came from out-of-sample data. Quantustik’s committed calibration figures come from exactly this kind of held-out evaluation — and we publish where the model still struggles rather than hiding it.
Overfit tools are seductive precisely because their backtests look perfect — that perfection is the warning sign, not the selling point. The one question that cuts through it: “Was this tested on data it had never seen?” If the answer is no, or unclear, the impressive numbers describe a past that already happened, not a future you can trade.
This lesson is investor education, not personalized advice. It teaches you to interrogate a backtest, not to trust one — no backtest, however honest, is a promise of future returns.
In-sample tests a model on the same data used to build and tune it — like grading a student on questions they were given the answers to. Out-of-sample tests it on data it never saw during tuning, which is the only fair predictor of real-world performance.
Real markets are noisy, so a track record with no losing periods usually means the model fit the random noise of the past (overfitting) rather than a repeatable pattern. That perfection rarely survives contact with data the model hasn’t seen.
“Was this tested on data the model had never seen?” If a tool can’t point to an out-of-sample or walk-forward result, its numbers describe the past it was fitted to, not a future you can trade on.
Investor education only — not investment advice, and never a promise of profit. Every investment can lose value.