How We Backtest: The Execution Rules That Keep Our Numbers Honest
How We Backtest: The Execution Rules That Keep Our Numbers Honest
Most backtests you see online are accidentally lying. Not maliciously — the code just makes quiet decisions that flatter the strategy: crediting a trade with a price move that happened before the signal, starting a moving average before it's fully formed, or dropping the exit day's return. Each one is small. Together they can turn a loser into a "proven system."
Here are the execution rules every StratIQ backtest follows, and why each one exists.
1. No look-ahead bias, ever
Rule: The signal is observed at bar i's close. The trade prints at bar i's close. The position starts earning on bar i+1.
Why: A signal computed from a day's closing price can't be traded at that day's open — or credited with that day's move. Crediting the entry bar's return to a new position is the single most common backtest bug, and it always flatters the strategy. We caught exactly this bug in our own QQQ golden-cross backtest and corrected the numbers publicly.
2. No phantom first-bar signals
Rule: Moving-average signals are evaluated only from the first bar where both SMAs are fully formed — and the cross requires a genuine transition (previous bar: short below long; current bar: short above long).
Why: On the very first bar where a 200-day SMA can be computed, there is no "previous" comparison — any signal there is an artifact of the data starting, not of the market. Our original golden-cross run bought on exactly such a phantom signal. Now the engine refuses to evaluate crosses before index 199 for a 50/200 pair.
3. Exit-day returns count
Rule: A position exited at bar i's close keeps bar i's return — it was held during that bar.
Why: This one cuts the other way: dropping the exit day's return understates a strategy. When we fixed this alongside the look-ahead bug, the corrected result was actually higher (+238.5% vs the flawed +228.5%) because the dropped exit-day returns outweighed the phantom entry-day gains. Honest accounting doesn't always hurt the strategy — but it has to be consistent.
4. Two implementations, one answer
Rule: Headline backtests are computed twice: once in our stratiq-backtest engine, once in a standalone script — both reading the same data file. All headline metrics must match.
Why: A bug in the engine and a bug in the test can cancel each other out. Two independent code paths reading identical data make that dramatically less likely.
5. Hand-verified signal dates
Rule: Every signal date in a published backtest is checked against hand-computed indicator values at full precision.
Why: Floating-point rounding near a crossover boundary can flip a signal by a day. Our 2020-05-19 golden-cross entry was a genuine knife-edge cross (50-day SMA 195.3198 vs 200-day 195.3152, prior day not a cross). If we hadn't verified it by hand, we'd never be sure the engine saw it right.
6. Disclosed data and costs
Rule: We state the data source (e.g., Yahoo Finance daily adjusted closes), the exact bar range, and that no commissions, slippage, or taxes are modeled.
Why: Adjusted closes vs raw closes, dividend treatment, and data vendor differences move results by tenths of a percent — enough to matter when comparing strategies. And a strategy that "works" before costs can easily fail after them. Real trading does worse than our numbers by roughly those costs.
The one-line version
A backtest you can't reproduce is a story, not a statistic. We publish the rules above so anyone can check our work — and we correct our own numbers in public when the rules catch our mistakes.
Not financial advice. These are research methods, not recommendations.