Summary
Running a backtest used to be slow and expensive. Now a single operator can test thousands of parameter combinations overnight. We argue that this abundance makes most backtest results less trustworthy, not more, and that the durable advantage in automated trading lies in two less glamorous places: running the same strategy code identically in every environment, and layered risk rules that can veto any order.
A backtest is a hypothesis, not evidence
A backtest answers a narrow question: how would this rule have done on this data, under these assumptions about costs and fills? It does not answer whether the rule works. Treating a good backtest as evidence of an edge is like treating a promising hypothesis as a result. The research step is generating hypotheses cheaply; the evidence comes later, from data the strategy has never seen and from markets it has to survive in.
Thousands of backtests, one survivor
Bailey, Borwein, López de Prado and Zhu showed that when enough strategy variations are tried, finding one with an impressive historical Sharpe ratio becomes likely by chance alone, and that such a strategy is likely to disappoint out of sample [1][2]. Their deflated Sharpe ratio adjusts a strategy's score for the number of trials it took to find it [3]. In academic finance, Harvey, Liu and Zhu argued that because hundreds of factors have been tested, a new one should clear a much higher bar of statistical significance than the usual standard [4].
The practical point: the count of backtests you ran is part of the result. A strategy found after five tries and one found after five thousand are not equally believable, even if their charts look identical.
Show the numbers
| Trials | Expected best Sharpe (zero skill) |
|---|---|
| 2 | 0.30 |
| 5 | 0.69 |
| 10 | 0.91 |
| 20 | 1.10 |
| 50 | 1.31 |
| 100 | 1.46 |
| 200 | 1.60 |
| 500 | 1.76 |
| 1,000 | 1.88 |
| 2,000 | 1.99 |
| 5,000 | 2.13 |
| 10,000 | 2.23 |
The numbers in Figure 1 are not from any particular market. They follow from arithmetic: with three years of daily data, a strategy with no edge still produces a Sharpe ratio estimate that wanders around zero, and the best of many such estimates drifts upward as the count grows.
Five questions for any backtest
- 01How many variations were tried before this one, including the ones nobody wrote down?
- 02What data was never touched during development, and how does the strategy do on it?
- 03What were the cost assumptions for fees, funding and slippage, and what happens to the result if they double?
- 04How many independent trades does the result rest on, and how much of the profit comes from the best few?
- 05Did the strategy run unchanged in paper trading against live data, and did the results match?
A backtest that can't answer these questions is a picture of a hypothesis. That can still be useful, as long as nobody mistakes it for evidence.
The counterargument
Not everyone agrees there is a crisis. Jensen, Kelly and Pedersen found that most published factors do replicate when analyzed carefully [5]. We take that seriously, with one distinction. Academic factors are tested on decades of data across many markets. A single operator sweeping thousands of parameter settings on a few years of crypto perpetuals data is in a much more dangerous position.
Why perpetuals raise the stakes
Crypto perpetual futures add hazards a simple backtest tends to skip. Funding payments change the cost of holding a position several times a day. Leverage turns small modeling errors into liquidations, and liquidations can cascade across a venue in minutes. Markets never close, so there is no daily pause in which to notice that something is wrong. And there are only a few years of history, which is exactly the setting where the overfitting arithmetic above bites hardest.
What disciplined backtesting looks like
- 01Record every trial, including the ones that failed, and report the count with the result.
- 02Keep a final stretch of data untouched until the very end, and use it once.
- 03Test with walk-forward or purged cross-validation, so information from the test period can't leak into training [8].
- 04Model fees, funding and slippage pessimistically. A strategy that only works with optimistic costs doesn't work.
One strategy, every environment
Backtests lie most where the backtest differs from reality: fills, fees, funding payments, latency and venue outages. The best defense we know is architectural. A strategy is written once and runs unchanged against a historical replay, a paper-trading sandbox and a live venue. Only the data source and the execution venue change. When results drift between environments, the drift is the finding, and it usually points to something the backtest assumed and the market didn't honor.
The code that says no
On August 1, 2012, a deployment error at Knight Capital left obsolete trading code active on one server. In about 45 minutes it sent millions of unintended orders, and the firm lost roughly $460 million, according to the SEC's findings [6]. No control between the strategy and the market stopped it. The SEC's market access rule exists to require exactly those pre-trade checks for broker-dealers [7].
In our design, every order passes layered limits before it leaves: exposure, leverage, drawdown and venue health. Any layer can veto. A person can pause or flatten everything at any time. These rules make no money on a good day. They decide whether there is a next day.
Every backtest you run makes the next one less believable. The durable edge is in the rules that refuse trades.
A person stays in charge
Automated trading systems are often described in terms of how little they need people. We think the opposite framing is safer. The operator can pause the system, flatten every position or tighten any limit at any moment, from anywhere, and those controls are tested as carefully as the strategy itself. A kill switch that has never been pulled in a drill is a hope, not a control.
The same principle covers changes. Knight Capital's loss began with a deployment. New strategy code and new limits should reach live trading only after running in paper trading against the same live data, and a change that alters risk limits should require a person's sign-off.
Why the gates must be independent
Risk checks that live inside the strategy share its bugs. If the strategy miscounts its position, a limit computed from that count is wrong in the same way. So the gates run separately from strategy code, keep their own record of positions reconciled against the venue, and fail closed: if their data is stale or missing, they block orders rather than let them through. A strategy can ask for a trade; it can never grant itself permission.
Where we might be wrong
- Vetoes have a cost. Limits that are too tight cut off the profitable tail that pays for the losses.
- Crypto perpetual markets trade around the clock and change regime quickly. Even honest backtests may describe a market that no longer exists.
- Perfect parity is impossible. A paper-trading sandbox never reproduces the market impact of real orders.
What we're measuring
- Trials per strategy: every backtest is recorded, and the count is reported with the result, alongside a deflated Sharpe ratio.
- Environment drift: the gap in returns, fills and costs between historical replay, paper trading and live trading for the same strategy.
- Veto accounting: how often each risk layer blocks an order, and what those orders would have earned or lost.
References
- [1]Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the AMS, 61(5). Link ↗
- [2]Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2017). The probability of backtest overfitting. Journal of Computational Finance, 20(4). Link ↗
- [3]Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting and non-normality. Journal of Portfolio Management, 40(5).
- [4]Harvey, C. R., Liu, Y., & Zhu, H. (2016). ...and the cross-section of expected returns. Review of Financial Studies, 29(1). Link ↗
- [5]Jensen, T. I., Kelly, B., & Pedersen, L. H. (2023). Is there a replication crisis in finance? Journal of Finance, 78(5).
- [6]U.S. Securities and Exchange Commission. (2013). In the Matter of Knight Capital Americas LLC. Release No. 70694.
- [7]U.S. Securities and Exchange Commission. (2010). Risk management controls for brokers or dealers with market access (Rule 15c3-5).
- [8]López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.