How We Backtest: The Execution Rules That Keep Our Numbers Honest

· methodology, backtesting, look-ahead bias

How We Backtest: The Execution Rules That Keep Our Numbers Honest

Most backtests you see online are accidentally lying. Not maliciously — the code just makes quiet decisions that flatter the strategy: crediting a trade with a price move that happened before the signal, starting a moving average before it's fully formed, or dropping the exit day's return. Each one is small. Together they can turn a loser into a "proven system."

Here are the execution rules every StratIQ backtest follows, and why each one exists.

1. No look-ahead bias, ever

Rule: The signal is observed at bar i's close. The trade prints at bar i's close. The position starts earning on bar i+1.

Why: A signal computed from a day's closing price can't be traded at that day's open — or credited with that day's move. Crediting the entry bar's return to a new position is the single most common backtest bug, and it always flatters the strategy. We caught exactly this bug in our own QQQ golden-cross backtest and corrected the numbers publicly.

2. No phantom first-bar signals

Rule: Moving-average signals are evaluated only from the first bar where both SMAs are fully formed — and the cross requires a genuine transition (previous bar: short below long; current bar: short above long).

Why: On the very first bar where a 200-day SMA can be computed, there is no "previous" comparison — any signal there is an artifact of the data starting, not of the market. Our original golden-cross run bought on exactly such a phantom signal. Now the engine refuses to evaluate crosses before index 199 for a 50/200 pair.

3. Exit-day returns count

Rule: A position exited at bar i's close keeps bar i's return — it was held during that bar.

Why: This one cuts the other way: dropping the exit day's return understates a strategy. When we fixed this alongside the look-ahead bug, the corrected result was actually higher (+238.5% vs the flawed +228.5%) because the dropped exit-day returns outweighed the phantom entry-day gains. Honest accounting doesn't always hurt the strategy — but it has to be consistent.

4. Two implementations, one answer

Rule: Headline backtests are computed twice: once in our stratiq-backtest engine, once in a standalone script — both reading the same data file. All headline metrics must match.

Why: A bug in the engine and a bug in the test can cancel each other out. Two independent code paths reading identical data make that dramatically less likely.

5. Hand-verified signal dates

Rule: Every signal date in a published backtest is checked against hand-computed indicator values at full precision.

Why: Floating-point rounding near a crossover boundary can flip a signal by a day. Our 2020-05-19 golden-cross entry was a genuine knife-edge cross (50-day SMA 195.3198 vs 200-day 195.3152, prior day not a cross). If we hadn't verified it by hand, we'd never be sure the engine saw it right.

6. Disclosed data and costs

Rule: We state the data source (e.g., Yahoo Finance daily adjusted closes), the exact bar range, and that no commissions, slippage, or taxes are modeled.

Why: Adjusted closes vs raw closes, dividend treatment, and data vendor differences move results by tenths of a percent — enough to matter when comparing strategies. And a strategy that "works" before costs can easily fail after them. Real trading does worse than our numbers by roughly those costs.

The one-line version

A backtest you can't reproduce is a story, not a statistic. We publish the rules above so anyone can check our work — and we correct our own numbers in public when the rules catch our mistakes.

Not financial advice. These are research methods, not recommendations.

← All posts