The standard advice — hold out the last 20% of your data and check the strategy there — is better than nothing and much weaker than most people believe. A single holdout is a single sample. Pass it and you have one piece of weak evidence; fail it and you will be tempted to adjust and try again, at which point the holdout has silently become training data.
Walk-forward analysis replaces that single split with a rolling procedure. Fit your parameters on a window of history, trade them forward on the period immediately following, then slide both windows and repeat. Concatenating the out-of-sample segments produces an equity curve made entirely of decisions taken with no knowledge of what came next. That curve, and not the in-sample one, is the honest estimate.
The choices that matter are window lengths and whether the fitting window rolls or expands. A short fitting window adapts quickly and overfits noise; a long one is stable and slow to notice a genuine regime change. For daily equity systems, two years fitting and three to six months forward is a reasonable starting point. Expanding windows suit strategies you believe are structural; rolling windows suit strategies you believe are adaptive. Pick on the basis of your thesis, then leave it alone.
The single most useful output is not the final return. It is walk-forward efficiency: out-of-sample performance divided by in-sample performance over the same segments. A ratio near 1.0 means the fitted parameters generalised. Around 0.5 is typical of a real but decaying edge. Below roughly 0.3 you have a curve-fitting procedure rather than a strategy, no matter how good the concatenated curve happens to look.
Watch parameter stability across segments too, because it often diagnoses the failure before the returns do. If the optimal lookback lands on 14, 15 and 13 across three consecutive fits, the surface has a plateau and the parameter means something. If it jumps from 9 to 47 to 22, the optimiser is chasing noise, and the fact that the concatenated out-of-sample curve happens to rise is luck you should not spend.
Be honest about how many times you have run the whole procedure. Walk-forward is not immune to the meta-overfitting problem: if you run twenty variants of a strategy through walk-forward and ship the one that passed, you have selected on the out-of-sample results and they are no longer out of sample. Budget your attempts in advance and record them, including the ones you abandoned.
Finally, resample. Monte Carlo over your trade sequence — shuffling order, or bootstrapping with replacement — turns one path into a distribution and answers the question that actually determines whether you can trade the thing: not "what did it return" but "how bad does the fifth-percentile drawdown get". Most strategies people abandon were abandoned inside a drawdown their own backtest had already predicted as ordinary.
Tomas Ekwueme
Quantitative Researcher
Writes for AlgoBeam on research. Every figure quoted above can be reproduced in the backtester with the same settings.