Statistics · Live Track Record

Ten thousand simulated years. Zero losses. That certainty is the flaw.

Grading a live track record with the statistics that institutions actually use — and finding out what a "can't-lose" result really means.

Home / Insights / Bootstrap & Deflated Sharpe
Statistics · Live Track Record  ·  August 2026  ·  6 min full read · 30 sec summary
TL;DR — 30 seconds
  • Bootstrap resampling of a live trading ledger produced ten thousand simulated years without a single losing one — a certainty that should raise suspicion, not confidence.
  • The record's annualised Sharpe of 5.1 passes a naive significance test; penalise it for the five strategy variants tried before this one, and confidence drops to 79%.
  • All three tools were blind in the same place: the sample contains no crisis. The honest output is a calendar date — the day the record actually earns formal significance.
5.1
annualised Sharpe over 56 live days
79.5%
probability the edge is real, once selection is penalised
~110 days
track length at which the record formally graduates

A live automated book had run for roughly two months. Every accounting-level check said the same thing: profitable, steady, boring. The natural next question was the one every allocator, every investment committee, and every desk eventually asks of a young track record: is this skill, or is this two lucky months? That question has real machinery behind it — and the machinery's answers are more interesting than a yes or a no.

Why this matters: most track records presented to investors are short, most are selected (you're shown the strategy that worked, not the ones quietly retired), and most are graded with tools that assume neither. The statistics below exist precisely to price those two facts in.

1Build the daily series honestly — boring days count

Fifty-six daily returns, including every day on which nothing happened.

Show the working

The raw material is unglamorous: every completed round-trip from the live ledger, netted of actual funding costs, aggregated into calendar-day returns on deployed capital. The tempting shortcut — and a common one in vendor material — is to compute statistics only over days with activity. That silently deletes the strategy's idle risk-free stretches and inflates every ratio built on top. The series used here includes all fifty-six days, zero-trade days as zeros.

2The naive verdict: impressive, barely significant

An annualised Sharpe of 5.08 — and a t-statistic of 1.99, which clears 95% confidence by a whisker.

Show the working

Annualised return of roughly 15% on 3% volatility gives a Sharpe ratio above 5 — a number that would headline any pitch deck. But a Sharpe ratio is an estimate, and estimates come with error bars set by track length. Fifty-six days of daily returns yield a t-statistic of 1.99: one-sided p of about 0.026. Statistically significant, yes — but only just, and only before anyone asks the two harder questions below. The same Sharpe over two years would be overwhelming evidence; over two months it is a promising whisper.

3Bootstrap: ten thousand parallel years

Resampling can only recombine the history you feed it — it cannot invent the bad day you haven't had yet.

Show the working

Bootstrap resampling draws random days from the observed series, with replacement, to assemble synthetic years — ten thousand of them here. The output distribution: median annualised return around 15%, a 90% confidence band of roughly +10% to +20%, worst simulated drawdown around 2%, and not one losing year in ten thousand.

Read it correctly

"Zero losing years in 10,000" is not evidence the strategy can't lose. It is evidence that the 56 observed days contain no day bad enough to sink a year, no matter how they are rearranged. The bootstrap faithfully answers "what do recombinations of my history look like" — it is structurally incapable of answering "what does the day I haven't seen look like." Every quiet sample bootstraps into invincibility.

Two further caveats belong in any honest use: day-by-day resampling assumes days are independent (real strategies cluster their good and bad stretches — block bootstrap variants exist for exactly this), and the input period here contained no market dislocation at all.

4Deflate it: pay for the shopping trip

Price in the five variants tried before this one, and "statistically significant" becomes 79.5% probable.

Show the working

The quiet fact behind most presented track records: the strategy you are shown is the survivor. This desk had honestly trialled about five parameter variants before settling on the current configuration. Selecting the best of five and then testing it as if it were the only one ever tried overstates the evidence — the best of five random strategies also looks good.

The Deflated Sharpe Ratio (Bailey & López de Prado) formalises the correction: it raises the significance hurdle to the Sharpe you'd expect the best of N tries to show by luck alone, and adjusts for the return distribution's skew and fat tails while it's at it.

Variants tried (N)Luck hurdle (Sharpe)Probability the edge is real
1 — pretending no selection0.098.2%
5 — the honest count3.0779.5%
104.0666.2%
204.9052.9% — a coin flip

Read the first row against the second: the difference between "98% proven" and "79% probable" is nothing about the strategy — it is purely whether the analyst admits how many things were tried. That single honesty parameter moves the verdict more than the data does. A Sharpe of five, over two months, best-of-five: a promising record that has not yet earned the word "significant."

5Notice what all three tools agree on

Three tools, one shared blind spot: the sample has no crisis in it.

Show the working

The bootstrap says "can't lose." The naive t-test says "significant." The deflator says "79.5%." They look like three verdicts; they are one verdict about the same hole, viewed from three angles. Every drop of certainty in these numbers is inherited from a 56-day window in which nothing violent happened — the same structural blind spot that made the Kelly criterion demand 3,000x leverage on this very book. Statistical machinery does not create information about the tail; it only makes the absence of that information easier to miss — or, used properly, easier to see.

6Convert the statistics into a calendar

Statistical patience has a date on it — and capital decisions can be scheduled against it.

Show the working

The Deflated Sharpe Ratio grows with the square root of track length. Holding performance at its current level, the arithmetic says this record crosses the 95% threshold — with the honest N = 5 penalty — at roughly 110 days of track: about two more months. That turns a vague "let's see more data" into an operational rule with a date: no scaling of capital before the record graduates; a scheduled re-grading when it does. The statistics' most useful output isn't a grade at all — it's a calendar.

What didn't happen: no capital was added on the strength of the "can't-lose" bootstrap, and no one quoted the Sharpe ratio without its track length attached. The review's output was a dated re-grading schedule and a documented honesty parameter — the N that most track-record presentations quietly omit.

The reusable checklist

Before accepting any track record — yours or a vendor's:

  • Does the daily return series include the idle days, or only the days with activity?
  • Is the Sharpe ratio quoted together with track length — and does its t-statistic actually clear significance?
  • What does a bootstrap of the record show — and does anyone acknowledge it can only recombine observed history, never extend it?
  • How many variants, funds, or configurations were tried before the one being shown? What does the Deflated Sharpe Ratio say at that honest N?
  • Does the sample period contain the market regime that actually threatens the strategy? If not, every ratio is a fair-weather ratio.
  • Is there a stated calendar date at which the record earns formal significance — and are capital decisions actually waiting for it?

Where this fits

This is the same discipline MOA applies to enterprise vendor claims under Independent Verification & Validation — extended to quantitative track records and performance claims. We review methodology and evidence; we don't sell or manage the strategies under review, and nothing here is personalised financial advice. No stake in the answer, no incentive to grade generously.

A track record, backtest, or performance claim that needs independent grading?

Methodology-first review before capital or client funds are committed — conflict-free.

Start a confidential conversation