Nobody sets out to fool themselves. The researchers who produce the most flattering backtests are usually the most careful ones, because care is what lets you iterate quickly, and rapid iteration is exactly the machine that manufactures false discoveries. The problem is structural, not moral. If you test two hundred variations of an idea against the same decade of price history, the best-looking result is a report on the noise in that decade rather than a report on the idea.
The first and largest leak is lookahead. It rarely appears as something obvious like using tomorrow’s close; it sneaks in through a restated fundamental, an index membership list that reflects today’s constituents, or an indicator that recalculates its own history when a new bar arrives. Our rule is that every series must be answerable to the question "what did this value read at the moment the bar closed?" If a data point cannot answer that, it is not allowed into a signal.
The second is survivorship. Testing a mean-reversion system on the current index membership tells you how mean reversion performed among companies that did not go bankrupt. Add the delisted tickers back and the same system typically loses somewhere between a fifth and a third of its edge, and the tail of the loss distribution gets considerably fatter. We include delisted names by default and we would rather you saw the uglier number.
The third is cost. A strategy that trades four hundred times a year on a mid-cap name is not paying commission alone; it is paying half the spread on entry and exit, plus impact that scales with how much of the bar’s range it demands. Flat per-trade slippage assumptions consistently flatter high-frequency systems. Slippage proportional to bar range is more pessimistic in fast markets, which is precisely when your signals cluster.
The fourth is the parameter search itself. If you optimise a lookback across forty values and report the best, you have not found a parameter, you have found the peak of a noise surface. The useful question is whether the neighbourhood is a plateau or a spike. A strategy whose profit factor is 1.8 at a 20-bar lookback and 1.7 at 18 and 22 has something. One that is 2.4 at 20 and 0.9 at 19 and 21 has nothing but luck.
The fifth is regime. Most of the equity history we all test on contains two enormous trending stretches and a handful of violent, short dislocations. A system fitted to that mix will look robust in aggregate and fall apart in the year it happens to face the wrong regime alone. Splitting your out-of-sample by volatility regime, rather than only by date, exposes this quickly and uncomfortably.
The sixth has no software fix. You will not trade the backtest. The backtest never skipped a signal because the last three lost, never doubled size to make back a bad week, and never went on holiday during the month that carried the year. Walk-forward windows, Monte Carlo resampling and honest cost modelling narrow the gap between the curve and reality. Closing the rest of it is a discipline problem, and it belongs to you.
Mara Vestergaard
Head of Quant Research
Writes for AlgoBeam on research. Every figure quoted above can be reproduced in the backtester with the same settings.