- Bootstrap resampling of a live trading ledger produced ten thousand simulated years without a single losing one — a certainty that should raise suspicion, not confidence.
- The record's annualised Sharpe of 5.1 passes a naive significance test; penalise it for the five strategy variants tried before this one, and confidence drops to 79%.
- All three tools were blind in the same place: the sample contains no crisis. The honest output is a calendar date — the day the record actually earns formal significance.
A live automated book had run for roughly two months. Every accounting-level check said the same thing: profitable, steady, boring. The natural next question was the one every allocator, every investment committee, and every desk eventually asks of a young track record: is this skill, or is this two lucky months? That question has real machinery behind it — and the machinery's answers are more interesting than a yes or a no.
Why this matters: most track records presented to investors are short, most are selected (you're shown the strategy that worked, not the ones quietly retired), and most are graded with tools that assume neither. The statistics below exist precisely to price those two facts in.
1Build the daily series honestly — boring days count
Fifty-six daily returns, including every day on which nothing happened.
Show the working
The raw material is unglamorous: every completed round-trip from the live ledger, netted of actual funding costs, aggregated into calendar-day returns on deployed capital. The tempting shortcut — and a common one in vendor material — is to compute statistics only over days with activity. That silently deletes the strategy's idle risk-free stretches and inflates every ratio built on top. The series used here includes all fifty-six days, zero-trade days as zeros.
2The naive verdict: impressive, barely significant
An annualised Sharpe of 5.08 — and a t-statistic of 1.99, which clears 95% confidence by a whisker.
Show the working
Annualised return of roughly 15% on 3% volatility gives a Sharpe ratio above 5 — a number that would headline any pitch deck. But a Sharpe ratio is an estimate, and estimates come with error bars set by track length. Fifty-six days of daily returns yield a t-statistic of 1.99: one-sided p of about 0.026. Statistically significant, yes — but only just, and only before anyone asks the two harder questions below. The same Sharpe over two years would be overwhelming evidence; over two months it is a promising whisper.
3Bootstrap: ten thousand parallel years
Resampling can only recombine the history you feed it — it cannot invent the bad day you haven't had yet.
Show the working
Bootstrap resampling draws random days from the observed series, with replacement, to assemble synthetic years — ten thousand of them here. The output distribution: median annualised return around 15%, a 90% confidence band of roughly +10% to +20%, worst simulated drawdown around 2%, and not one losing year in ten thousand.
"Zero losing years in 10,000" is not evidence the strategy can't lose. It is evidence that the 56 observed days contain no day bad enough to sink a year, no matter how they are rearranged. The bootstrap faithfully answers "what do recombinations of my history look like" — it is structurally incapable of answering "what does the day I haven't seen look like." Every quiet sample bootstraps into invincibility.
Two further caveats belong in any honest use: day-by-day resampling assumes days are independent (real strategies cluster their good and bad stretches — block bootstrap variants exist for exactly this), and the input period here contained no market dislocation at all.
4Deflate it: pay for the shopping trip
Price in the five variants tried before this one, and "statistically significant" becomes 79.5% probable.
Show the working
The quiet fact behind most presented track records: the strategy you are shown is the survivor. This desk had honestly trialled about five parameter variants before settling on the current configuration. Selecting the best of five and then testing it as if it were the only one ever tried overstates the evidence — the best of five random strategies also looks good.
The Deflated Sharpe Ratio (Bailey & López de Prado) formalises the correction: it raises the significance hurdle to the Sharpe you'd expect the best of N tries to show by luck alone, and adjusts for the return distribution's skew and fat tails while it's at it.
| Variants tried (N) | Luck hurdle (Sharpe) | Probability the edge is real |
|---|---|---|
| 1 — pretending no selection | 0.0 | 98.2% |
| 5 — the honest count | 3.07 | 79.5% |
| 10 | 4.06 | 66.2% |
| 20 | 4.90 | 52.9% — a coin flip |
Read the first row against the second: the difference between "98% proven" and "79% probable" is nothing about the strategy — it is purely whether the analyst admits how many things were tried. That single honesty parameter moves the verdict more than the data does. A Sharpe of five, over two months, best-of-five: a promising record that has not yet earned the word "significant."
5Notice what all three tools agree on
Three tools, one shared blind spot: the sample has no crisis in it.
Show the working
The bootstrap says "can't lose." The naive t-test says "significant." The deflator says "79.5%." They look like three verdicts; they are one verdict about the same hole, viewed from three angles. Every drop of certainty in these numbers is inherited from a 56-day window in which nothing violent happened — the same structural blind spot that made the Kelly criterion demand 3,000x leverage on this very book. Statistical machinery does not create information about the tail; it only makes the absence of that information easier to miss — or, used properly, easier to see.
6Convert the statistics into a calendar
Statistical patience has a date on it — and capital decisions can be scheduled against it.
Show the working
The Deflated Sharpe Ratio grows with the square root of track length. Holding performance at its current level, the arithmetic says this record crosses the 95% threshold — with the honest N = 5 penalty — at roughly 110 days of track: about two more months. That turns a vague "let's see more data" into an operational rule with a date: no scaling of capital before the record graduates; a scheduled re-grading when it does. The statistics' most useful output isn't a grade at all — it's a calendar.
What didn't happen: no capital was added on the strength of the "can't-lose" bootstrap, and no one quoted the Sharpe ratio without its track length attached. The review's output was a dated re-grading schedule and a documented honesty parameter — the N that most track-record presentations quietly omit.
The reusable checklist
Before accepting any track record — yours or a vendor's:
- Does the daily return series include the idle days, or only the days with activity?
- Is the Sharpe ratio quoted together with track length — and does its t-statistic actually clear significance?
- What does a bootstrap of the record show — and does anyone acknowledge it can only recombine observed history, never extend it?
- How many variants, funds, or configurations were tried before the one being shown? What does the Deflated Sharpe Ratio say at that honest N?
- Does the sample period contain the market regime that actually threatens the strategy? If not, every ratio is a fair-weather ratio.
- Is there a stated calendar date at which the record earns formal significance — and are capital decisions actually waiting for it?
Where this fits
This is the same discipline MOA applies to enterprise vendor claims under Independent Verification & Validation — extended to quantitative track records and performance claims. We review methodology and evidence; we don't sell or manage the strategies under review, and nothing here is personalised financial advice. No stake in the answer, no incentive to grade generously.