Research/Crypto Research/Position paper

The most valuable code in a trading system is the code that says no

Modern tools make thousands of backtests cheap. Statistics says that's exactly the problem.

Supersafe Research · October 6, 2026 · 8 min read

Summary

Running a backtest used to be slow and expensive. Now a single operator can test thousands of parameter combinations overnight. We argue that this abundance makes most backtest results less trustworthy, not more, and that the durable advantage in automated trading lies in two less glamorous places: running the same strategy code identically in every environment, and layered risk rules that can veto any order.

A backtest is a hypothesis, not evidence

A backtest answers a narrow question: how would this rule have done on this data, under these assumptions about costs and fills? It does not answer whether the rule works. Treating a good backtest as evidence of an edge is like treating a promising hypothesis as a result. The research step is generating hypotheses cheaply; the evidence comes later, from data the strategy has never seen and from markets it has to survive in.

Thousands of backtests, one survivor

Bailey, Borwein, López de Prado and Zhu showed that when enough strategy variations are tried, finding one with an impressive historical Sharpe ratio becomes likely by chance alone, and that such a strategy is likely to disappoint out of sample [1][2]. Their deflated Sharpe ratio adjusts a strategy's score for the number of trials it took to find it [3]. In academic finance, Harvey, Liu and Zhu argued that because hundreds of factors have been tested, a new one should clear a much higher bar of statistical significance than the usual standard [4].

The practical point: the count of backtests you ran is part of the result. A strategy found after five tries and one found after five thousand are not equally believable, even if their charts look identical.

BEST BACKTEST SHARPE FROM PURE LUCK · 3 YEARS OF DATA · ZERO TRUE SKILL0.00.51.01.52.02.51101001,00010,000STRATEGY VARIATIONS TRIED (LOG SCALE)10 trials: expected best Sharpe ≈ 0.910.9100 trials: expected best Sharpe ≈ 1.461.51,000 trials: expected best Sharpe ≈ 1.881.910,000 trials: expected best Sharpe ≈ 2.232.2Expected best of N zero-skill backtests (Bailey & López de Prado). Computed, not measured.
Show the numbers
TrialsExpected best Sharpe (zero skill)
20.30
50.69
100.91
201.10
501.31
1001.46
2001.60
5001.76
1,0001.88
2,0001.99
5,0002.13
10,0002.23
Figure 1 With no real edge at all, the best of many backtests still looks good. With three years of data, trying 1,000 variations should produce a best Sharpe ratio near 1.9 by luck alone.

The numbers in Figure 1 are not from any particular market. They follow from arithmetic: with three years of daily data, a strategy with no edge still produces a Sharpe ratio estimate that wanders around zero, and the best of many such estimates drifts upward as the count grows.

Five questions for any backtest

  1. 01How many variations were tried before this one, including the ones nobody wrote down?
  2. 02What data was never touched during development, and how does the strategy do on it?
  3. 03What were the cost assumptions for fees, funding and slippage, and what happens to the result if they double?
  4. 04How many independent trades does the result rest on, and how much of the profit comes from the best few?
  5. 05Did the strategy run unchanged in paper trading against live data, and did the results match?

A backtest that can't answer these questions is a picture of a hypothesis. That can still be useful, as long as nobody mistakes it for evidence.

The counterargument

Not everyone agrees there is a crisis. Jensen, Kelly and Pedersen found that most published factors do replicate when analyzed carefully [5]. We take that seriously, with one distinction. Academic factors are tested on decades of data across many markets. A single operator sweeping thousands of parameter settings on a few years of crypto perpetuals data is in a much more dangerous position.

Why perpetuals raise the stakes

Crypto perpetual futures add hazards a simple backtest tends to skip. Funding payments change the cost of holding a position several times a day. Leverage turns small modeling errors into liquidations, and liquidations can cascade across a venue in minutes. Markets never close, so there is no daily pause in which to notice that something is wrong. And there are only a few years of history, which is exactly the setting where the overfitting arithmetic above bites hardest.

What disciplined backtesting looks like

  1. 01Record every trial, including the ones that failed, and report the count with the result.
  2. 02Keep a final stretch of data untouched until the very end, and use it once.
  3. 03Test with walk-forward or purged cross-validation, so information from the test period can't leak into training [8].
  4. 04Model fees, funding and slippage pessimistically. A strategy that only works with optimistic costs doesn't work.

One strategy, every environment

Backtests lie most where the backtest differs from reality: fills, fees, funding payments, latency and venue outages. The best defense we know is architectural. A strategy is written once and runs unchanged against a historical replay, a paper-trading sandbox and a live venue. Only the data source and the execution venue change. When results drift between environments, the drift is the finding, and it usually points to something the backtest assumed and the market didn't honor.

One strategysame code, every environmentONE EVENT INTERFACE · CANDLES, TICKS, FUNDING, FILLSHistorical replayrecorded market dataPaper tradinglive data, simulated fillsLive venuereal orders, real moneyDRIFTDRIFTWhen results differ between environments, the difference is the finding.
Figure 2 One strategy, unchanged, runs against historical replay, paper trading and a live venue through a single event interface. Differences between them are measured as drift.

The code that says no

On August 1, 2012, a deployment error at Knight Capital left obsolete trading code active on one server. In about 45 minutes it sent millions of unintended orders, and the firm lost roughly $460 million, according to the SEC's findings [6]. No control between the strategy and the market stopped it. The SEC's market access rule exists to require exactly those pre-trade checks for broker-dealers [7].

In our design, every order passes layered limits before it leaves: exposure, leverage, drawdown and venue health. Any layer can veto. A person can pause or flatten everything at any time. These rules make no money on a good day. They decide whether there is a next day.

Every backtest you run makes the next one less believable. The durable edge is in the rules that refuse trades.
EVERY ORDER, EVERY TIMEOrderExposureVETOLeverageVETODrawdownVETOVenue healthVETOVenuePERSON: PAUSE OR FLATTEN AT ANY TIMEKnight Capital, 2012:no gate like these stood between faulty code and the market.About 45 minutes; losses of roughly $460 million (SEC findings).
Figure 3 Every order passes layered risk gates before it reaches a venue, and any gate can veto it. A person can pause or flatten everything at any time.

A person stays in charge

Automated trading systems are often described in terms of how little they need people. We think the opposite framing is safer. The operator can pause the system, flatten every position or tighten any limit at any moment, from anywhere, and those controls are tested as carefully as the strategy itself. A kill switch that has never been pulled in a drill is a hope, not a control.

The same principle covers changes. Knight Capital's loss began with a deployment. New strategy code and new limits should reach live trading only after running in paper trading against the same live data, and a change that alters risk limits should require a person's sign-off.

Why the gates must be independent

Risk checks that live inside the strategy share its bugs. If the strategy miscounts its position, a limit computed from that count is wrong in the same way. So the gates run separately from strategy code, keep their own record of positions reconciled against the venue, and fail closed: if their data is stale or missing, they block orders rather than let them through. A strategy can ask for a trade; it can never grant itself permission.

Where we might be wrong

  • Vetoes have a cost. Limits that are too tight cut off the profitable tail that pays for the losses.
  • Crypto perpetual markets trade around the clock and change regime quickly. Even honest backtests may describe a market that no longer exists.
  • Perfect parity is impossible. A paper-trading sandbox never reproduces the market impact of real orders.

What we're measuring

  • Trials per strategy: every backtest is recorded, and the count is reported with the result, alongside a deflated Sharpe ratio.
  • Environment drift: the gap in returns, fills and costs between historical replay, paper trading and live trading for the same strategy.
  • Veto accounting: how often each risk layer blocks an order, and what those orders would have earned or lost.

References

  1. [1]Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-mathematics and financial charlatanism: The effects of backtest overfitting on out-of-sample performance. Notices of the AMS, 61(5). Link ↗
  2. [2]Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2017). The probability of backtest overfitting. Journal of Computational Finance, 20(4). Link ↗
  3. [3]Bailey, D. H., & López de Prado, M. (2014). The deflated Sharpe ratio: Correcting for selection bias, backtest overfitting and non-normality. Journal of Portfolio Management, 40(5).
  4. [4]Harvey, C. R., Liu, Y., & Zhu, H. (2016). ...and the cross-section of expected returns. Review of Financial Studies, 29(1). Link ↗
  5. [5]Jensen, T. I., Kelly, B., & Pedersen, L. H. (2023). Is there a replication crisis in finance? Journal of Finance, 78(5).
  6. [6]U.S. Securities and Exchange Commission. (2013). In the Matter of Knight Capital Americas LLC. Release No. 70694.
  7. [7]U.S. Securities and Exchange Commission. (2010). Risk management controls for brokers or dealers with market access (Rule 15c3-5).
  8. [8]López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.