Backtesting measures how a trading strategy would have performed on historical data. Forward testing applies its rules to prices arriving after those rules were fixed. Strategy validation uses both, alongside checks on data, costs and execution, to assess whether the results deserve further consideration.
The aim is not to produce the smoothest equity curve. It is to find reasons a strategy might fail before committing money. A profitable simulation is evidence to examine, not permission to trade.
Backtesting, Walk Forward Testing and Forward Testing
These methods answer different questions. Keeping them separate prevents a common mistake: treating several views of the same historical data as independent confirmation.
| Method | What it does | Main constraint |
|---|---|---|
| Historical backtest | Applies trading rules to past market data. | Results depend on data quality, assumptions and how the rules were selected. |
| Walk forward test | Repeatedly develops or calibrates rules on an earlier window, then tests the following window. | It remains a historical simulation and can still be overfitted. |
| Forward paper test | Records simulated trades as new market data arrives. | Simulated fills and behaviour may differ from actual trading. |
| Small live trial | Places actual orders under a restricted risk budget. | Money is at risk, and a short record cannot establish durable profitability. |
A historical period excluded from development is an out of sample test. It is not the same as watching a strategy operate prospectively. The CFA Institute material on backtesting and simulation covers the rolling window approach and the weaknesses of relying on historical returns.
Define the Rules Before Measuring Performance
Start with a testable proposition rather than an indicator collection. “Buy strong markets” leaves too much room for hindsight. A usable rule identifies the market, signal calculation, decision time, order type and exit conditions.
Write down position sizing, maximum simultaneous positions, trading hours and any conditions that prohibit entry. Decide how to handle missing prices, gaps and conflicting signals. These instructions belong in a written trading plan, alongside the circumstances that would stop the test.
For a hypothetical breakout strategy, distinguish between buying when price crosses yesterday’s high and buying after a session closes above it. Those are different strategies, with different information and execution requirements.
Manual testing also needs fixed rules. Hide future bars during chart replay and record decisions before revealing the outcome. For automated testing, check a sample of trades by hand against the written specification. Code that runs without errors can still implement the wrong strategy.
Build a Backtest That Could Have Happened
Use information available at the time
Look ahead bias occurs when a historical decision uses information that was not yet available. Examples include using revised economic figures instead of the original release or letting an indicator access a completed candle before it closed.
Survivorship bias arises when the dataset excludes securities that disappeared. Testing an old share strategy using only today’s surviving companies can remove failures from the record. Both problems are addressed in the CFA Institute Research Foundation guide to investment model validation.
Record the data provider, download date, timezone and treatment of corporate actions. For shares, check splits, dividends and delistings. For futures, document contract rolls. For currency strategies, establish whether prices represent bids, offers or midpoints. Do not assume every historical price could have been traded.
Audit the order assumptions
If a signal needs a candle’s closing price, do not automatically award an execution at that same price. Model the first feasible order opportunity after the information becomes available.
A candle containing both a stop price and a profit target does not reveal which was reached first. Use finer data where available or test a conservative ordering assumption. A price touching a limit order also does not prove the order would have filled.
Backtesting engines make choices about these situations. TradingView’s strategy simulation documentation details its historical price path assumptions and order timing. Whatever software you use, inspect those settings rather than treating the default result as a market fact.
Keep Development Separate From Validation
Divide the research process into development, validation and a final untouched test. Use development data to build the rules. Use validation data to compare candidates. Reserve the final test for the version selected before seeing its results.
Keep the sequence chronological when assessing how a strategy would have operated through time. Any fitted input, such as a prediction model or scaling calculation, must use only information available before the simulated decision. Where training outcomes extend into the test period, exclude the overlapping observations.
A walk forward design might calibrate on three years, test the next six months, then advance six months and repeat. Those durations are illustrative, not recommended defaults. Combine the subsequent test windows into one performance record and include the costs of any resulting portfolio changes.
Repeated experimentation creates another problem. Trying hundreds of combinations can produce an attractive winner by chance. The research paper The Probability of Backtest Overfitting examines why selecting the best historical result can lead to disappointing performance outside the selection sample.
Keep a record of rejected variants, not just the winner. If you inspect the final test, change the strategy and run it again, that period has become development data. Renaming it “validation” does not restore its independence. An untouched test helps, but it does not erase the effects of a large search.
Measure Returns After Costs and Alongside Risk
Include commissions, bid and offer spreads, slippage, financing and borrowing charges where applicable. Match them to the instrument and holding period. Avoid double counting: a simulation that buys at the offer and sells at the bid already includes that spread effect.
Consider a fictional set of 100 trades with equal position sizes. There are 45 winners averaging £120 and 55 losers averaging £80, before costs:
Average gross result = (0.45 × £120) − (0.55 × £80) = £10 per trade.
If the average total cost of opening and closing each trade is £12, the average net result becomes a £2 loss. Gross profit of £1,000 becomes a £200 loss across the sample. Neither the win rate nor the gross profit tells the full story.
Alongside net returns, report maximum drawdown, recovery time, trading frequency and time spent holding positions. Drawdown should reflect account equity, including open positions, rather than only completed trades. The separate guide to measuring trading results after costs covers performance accounting in more detail.
Compare the strategy with a sensible alternative over the same dates, such as cash or a relevant passive holding. Account for differences in exposure and risk; raw return alone is not a fair comparison. Also show how much profit comes from the largest winners, one instrument or one short period.
Stress Test the Result, Not Just the Settings
Before running sensitivity checks, decide which assumptions deserve scrutiny. Otherwise the stress test can become another search for flattering settings.
Useful checks include higher trading costs, delayed entries, fewer limit order fills and modest changes to lookback periods or exit distances. Test neighbouring settings rather than searching only for the single best value. If a 20 period rule works but nearby values collapse, investigate the dependence rather than celebrating the precision.
Inspect different market conditions and calendar periods. A trend strategy need not profit in every sideways market, but its losses there must fit the proposed risk budget. Removing its largest winning trades can show concentration risk, though it is a diagnostic exercise rather than a requirement that every strategy pass.
Do not treat a fixed trade count as a certificate of validity. A hundred positions opened during one market event do not provide the same breadth of evidence as observations spread across changing conditions. Ask how much independent information the sample contains, not just how many rows appear in the spreadsheet.
Record uncertainty explicitly. If removing one month reverses the result, or the strategy barely covers plausible costs, “inconclusive” is a useful finding. A test does not owe you a yes.
Run a Forward Test Without Moving the Goalposts
Freeze the strategy version before starting. Use the intended instruments, trading hours and position sizing rules. Set a review schedule in advance rather than ending the test immediately after a strong run.
Record every qualifying signal, including missed trades, rejected orders and occasions when the platform or data feed failed. Separate the strategy’s theoretical result from the result produced by the actual workflow. Otherwise operational mistakes disappear from the report while remaining very real in practice.
Paper trading avoids putting capital at risk, but it cannot fully reproduce actual execution or the pressure of losing money. These weaknesses are identified in the NFA notice on hypothetical performance results. A profitable paper record therefore does not establish that the same returns were achievable with real orders.
Compare forward results with the backtest at the trade level. Are signals arriving at the expected time? Are spreads wider? Are orders missing fills? Are overnight charges different? Investigate the mechanism behind a mismatch before deciding that the strategy has stopped working.
Choose the test duration around opportunity frequency and the conditions observed. Three months producing six trades answers a different question from three months producing hundreds. The guide to what demo accounts can and cannot teach covers the practical boundaries of simulated trading.
Decide Whether to Reject, Continue or Trial
Set decision criteria before viewing the final results. Use numerical limits suited to the proposed strategy and affordable risk, rather than borrowing a universal Sharpe ratio, win rate or drawdown threshold.
- Reject or redesign: results depend on unavailable information, implausible fills or costs that cannot be achieved.
- Continue testing: the process is sound, but the sample is too narrow or the net advantage too uncertain.
- Consider a restricted live trial: historical and forward evidence meet the predefined criteria, operational checks pass and the potential loss is affordable.
A live trial is optional, not the reward for completing a backtest. If undertaken, define a loss budget and stop conditions before placing orders. Do not increase exposure simply to recover an early loss; loss chasing and knowing when to stop deserve their own controls.
Keep the data snapshot, rules, software version, assumptions and full trade record so the result can be reproduced. Reassess when execution costs, market behaviour or implementation changes. Validation supports a decision under stated assumptions. It never turns an uncertain trading strategy into a guarantee.