The Backtest Claimed +60% in 11 Months. The Real Number Is Closer to 5%.
A published gold strategy came with a spectacular backtest. We retested the same rules on seven years of tick data with real costs, swept the complete parameter grid, and gave the survivors one shot at four unseen years. The edge is real. It is also five times smaller than advertised, and half of it lives in a single year.

Larry Williams' volatility breakout is a classic. When price moves a meaningful fraction of yesterday's range away from today's open, that is real order flow rather than noise, so you go with it. An MQL5 article implemented it for gold and reported the kind of backtest that sells EAs: +60% in 11 months, "steady progression," no extreme drawdowns.
The fine print is easy to miss. The test covered January to November 2025 only, one hand-picked year on daily bars in one of gold's great bull runs, and it charged zero commission.
So we ran the test the claim deserved: the complete parameter grid on three years of tick data, then a single shot at four more years the rules had never seen. The answer is more interesting than either "it works" or "it doesn't."
Our version of the test
Same rules, implemented faithfully. Buy when price closes above
day open + K x yesterday's range (mirror short below), stop at a fraction
of yesterday's range, take-profit at a reward multiple, one trade per day,
flat before the next day's levels.
Instead of one lucky year, we swept every combination of breakout multiple (0.3 to 0.8), stop multiple (0.3 to 0.7) and reward ratio (2 to 5). Seventy two cells, 2019 to 2022, real tick data, commission and swap included. Not a genetic sample that might miss pockets. Every cell of the space.
Sixty two of the 72 made money in sample. The best cells cluster in one neighbourhood rather than sitting alone on a spike, which is the shape you want: a breakout multiple of 0.5, a wide stop at 0.7 of yesterday's range, and a reward target of 3 to 5. Confirmed on real ticks, the leader returned +$11,964 on $10,000 across 479 trades with a 7.9% drawdown.
That is a great in-sample result, and in-sample results are cheap. The whole point of the next step is that we only get one.
Four years the rules had never seen
We froze three candidates, wrote the parameters down, and ran each one once over 2022 to 2026. No tuning afterwards. That window is now spent.
| Candidate | In-sample 2019-2022 | Out-of-sample 2022-2026 | Drawdown |
|---|---|---|---|
| Breakout 0.5 / stop 0.7 / RR3 | +$11,964 (PF 1.45) | +$1,654 (PF 1.07, 675 trades) | 15.8% |
| Breakout 0.5 / stop 0.7 / RR5 | +$12,423 (PF 1.47) | +$1,781 (PF 1.07, 675 trades) | 15.9% |
| Breakout 0.7 / stop 0.7 / RR3 | +$6,497 (PF 1.39) | +$3,502 (PF 1.20, 502 trades) | 23.5% |
All three finished profitable. All three survive a cost stress test that doubles commission. The edge is real, and if you had traded the best of them through 2022 to 2025 you would have made money.
Best candidate out-of-sample: two years underwater before 2025 carries the result
You would also have been miserable doing it. The strongest candidate spent its first year underwater (2022: -$299), crawled through 2023, and made most of its four-year profit in 2025 alone: -299, +74, +880, +2,847. The two conservative candidates actually lost money in 2023. A 23.5% drawdown on a strategy earning under 8% a year is a punishing ratio to live with.
Why it fails our gates
Our rule is that a strategy must keep at least half its in-sample return when it meets unseen data. These keep 13% to 43%. In-sample the strategy paid 18% to 31% a year; out of sample it pays 3.9% to 7.8%. Drawdowns doubled or tripled. So it is rejected, and it joins the growing pile of strategies that are real but too small to own.
The nuance worth sitting with: the profit factor held up far better than the return did, retaining 73% to 86%. The signal did not stop working. It got smaller and choppier, because the in-sample window was COVID plus a gold bull market and the years that followed were ordinary. Most published backtests are measuring the regime and calling it an edge.
How much of that was luck
"All three finished profitable" is a sentence about one realized path. The honest follow-up is to ask how easily it could have gone the other way, so we resampled each candidate's actual trade list ten thousand times.
Each run reshuffles the 283 out-of-sample trades, drawing them with replacement, and totals the result. The blue line is what actually happened; the shaded band left of zero is every run that ended in the red.
The answer splits the three apart. The aggressive candidate holds up: it loses money in 5.7% of resamples, and its profit sits in a range of roughly -$890 to +$7,885. The two conservative candidates do not. They lose money in about one resample in four, and their 2.5th-percentile profit factor is 0.88. Their +$1,654 and +$1,781 are not really results; they are the middle of a wide distribution that comfortably includes losing. If you only read the table above, you would never know that two of those three green numbers were coin flips.
The drawdown goes the other way, and in the strategy's favour. Shuffling the order of the same trades puts the typical worst drawdown near 12%, with 19% at the 95th percentile. The 23.5% we actually observed sits at the 99th percentile of orderings, so this particular sequence was an unlucky one rather than a flattering one. Treat that gently, though: shuffling assumes one trade tells you nothing about the next, and this strategy earned most of its money in a single year. Destroying that clustering is exactly what makes the reshuffled drawdowns look tamer than the life you would have lived.
Did it beat doing nothing clever?
This is the question that settles it. Over the same four years, simply holding gold returned 24.0% a year with a 20.4% drawdown, a MAR of 1.18. A mindless trend follower that goes long whenever the past year was up scored 1.24. The strategy scored 0.33.
Holding gold paid three times the return and drew down less. The complexity bought nothing: on a risk-adjusted basis the strategy is about three and a half times worse than the thing you could have done by falling asleep. That is a harder verdict than the retention gates deliver on their own, and it is the one worth remembering, because it applies to a lot more than this strategy. A gold system tested through a gold bull market has to beat gold, not zero.
A data convention worth knowing about
One market-structure detail matters for anyone testing daily-level systems on UTC data: the Sunday-evening open creates tiny "Sunday" daily bars (a $3 to $6 range versus $11 to $60 for real trading days). Any "yesterday's range" logic must treat Mondays specially, using Friday's full range, and never trade the Sunday stub itself, or roughly 20% of trades run on corrupted levels. Our implementation handles this, and every number in this post reflects it.
A second one, specific to gold: the market closes for the weekend and takes a daily break, so a strategy with a 23-hour time exit cannot always exit in 23 hours. About a fifth of these trades are held over a weekend and pay three nights of swap for the privilege. That cost is inside every figure above.
So what was the +60% worth?
Not nothing, which is the surprise. The rules do contain a real gold breakout edge. But the honest version of the claim is "under 8% a year with a 23% drawdown, most of it earned in one good year," not "+60% in eleven months." The gap between those two sentences is what a single hand-picked period, zero costs, and no holdout window buy you.
A note on the breakout family
This is our second data point on breakouts. The Asian-session range breakout, a session-anchored cousin of this idea, showed a genuine multi-year edge on USD/JPY while failing on EUR/USD. This daily-open volatility expansion is real on gold but too small to clear our bar. The family label tells you very little; what matters is the specific mechanism on the specific instrument, and how much of the result survives contact with costs and unseen years.
All tested strategies, winners and losers, live on the results page.
Past performance is not indicative of future results. These are backtests with realistic cost assumptions, not live trading records.
Run the numbers yourself
The free calculators behind the sizing and robustness checks in this test. No signup.