The Kelly criterion is one of the few pieces of trading mathematics with an actual theorem behind it: bet the fraction of your bankroll that maximises the expected logarithm of wealth, and in the long run you compound faster than any other strategy. It is elegant, it is provably optimal under its assumptions, and it is routinely either worshipped or dismissed — usually without anyone running it against real numbers.
We ran it against real numbers: a few hundred live fills from an automated market-neutral book — actual executions, actual funding costs, no simulation. What the formula returned, and the specific ways in which it was wrong, turned out to be far more instructive than the formula itself.
Why this matters: position sizing is where most quantitative blow-ups actually happen — not in the signal, but in how much was bet on it. A sizing rule that looks rigorous because it has a formula attached is more dangerous than an honest heuristic, because the formula hides which of its inputs is doing all the work.
1Run the naive calculation honestly
From the live ledger, each completed round-trip's net return — after real, per-trade funding costs at current borrow rates — was expressed as a fraction of position size. The distribution was unglamorous and healthy: roughly a 75% win rate, a mean of about half a basis point per trade, a worst single trade of a few basis points, hundreds of trades per year.
Feed that distribution into the standard continuous Kelly approximation (mean over variance) and it returns an optimal leverage of roughly three thousand times.
A formula that answers "3,000x" is not telling you to lever 3,000x. It is telling you that the distribution you fed it contains almost no risk — and since the strategy demonstrably does carry risk, the sample must be missing the part of the distribution where that risk lives. The absurd output is the most useful thing Kelly produces: it is a data-completeness alarm, not a bet size.
The reason the number is huge is mechanical: in calm markets this class of strategy loses rarely and loses small, so the sample variance is tiny, and Kelly divides by it. The catastrophic scenario — a large, fast market dislocation that runs through every protective layer at once — has simply never appeared in the sample. Not because it can't happen; because it hasn't happened during the measurement window.
2Put the tail back in — and watch the answer collapse
The fix is straightforward to write down: augment the observed distribution with a crisis scenario — one severe dislocation every few years, with some assumed per-position adverse move A — and recompute expected log growth as a function of leverage. The table below is the actual output, with the crisis frequency held fixed and only the severity assumption varied:
| Leverage | No tail (as sampled) | Tail A = 4% | Tail A = 8% | Tail A = 12% |
|---|---|---|---|---|
| 1x | +1.8% | +0.4% | −1.0% | −2.5% |
| 3x | +5.3% | +1.1% | −3.8% | −9.5% |
| 5x | +8.9% | +1.4% | −8.2% | −21.7% |
| 8x | +14.2% | +1.3% | −19.9% | −93.1% |
| 10x | +17.7% | +0.7% | −35.9% | ruin |
| 20x | +35.4% | −18.2% | ruin | ruin |
| 3,000x | "optimal" | ruin | ruin | ruin |
Read across any row and the story is stark. Under a 4% tail assumption, optimal leverage sits somewhere in the 5–10x range. At 8%, the optimum collapses to unlevered. At 12%, every leveraged variant has negative long-run growth and the higher tiers are outright ruinous. The optimal answer moves by three orders of magnitude across tail assumptions that the historical record cannot distinguish between — the sample contains approximately one crisis-class event, which is a sample size of one.
The core finding: Kelly is a telescope pointed at exactly the part of the distribution you cannot see. Its decisive input — tail severity at your chosen leverage — is the single quantity history refuses to hand over. The formula isn't wrong; it's answering a question you don't have the data to ask.
One validating detail is worth recording: the no-tail model's predicted growth at the leverage the book actually runs matched the book's live annualised return to within a percentage point. The model reproduces reality where reality has been observed. The entire disagreement is about the part that hasn't.
3Name the other broken assumptions — there are three more
Tail blindness is the headline failure, but Kelly's premises break in three quieter ways for this class of strategy, and each one matters for anyone tempted to apply the formula to a market-making or liquidity-provision book:
- Returns are not independent draws. The book's profitability is regime-driven — long stretches of small, similar outcomes punctuated by rare windows where most of the annual profit arrives. Kelly's optimality proof assumes i.i.d. bets; a regime-switching process needs regime-conditional sizing, not one global fraction.
- Bet size changes the game itself. Kelly assumes a price-taker whose stake doesn't alter the odds. A resting-order strategy's fill rate is a function of its own size — double the order and you don't double the profit, you halve the queue position. Long before Kelly's fraction binds, order-book capacity binds. In calm regimes, the effective constraint on size was found to be market capacity at roughly one-thirtieth of what unconstrained Kelly implied.
- Waiting isn't free. The classic setup has you betting discrete rounds with capital idle between them. A leveraged carry book pays financing continuously while positions wait — the cost side scales with time-in-market, not bet count. At current rates this alone reshapes which trades are worth taking, independent of any sizing formula.
4Notice what the desk already does — Kelly's skeleton without Kelly's estimate
Here is the part that surprised the review: examined structurally, the book's existing sizing discipline turns out to implement almost everything the Kelly framework actually teaches, while refusing the one thing it cannot deliver — a point estimate built on an unobservable input.
- Fixed-fraction compounding. Position notional is recomputed from live equity on every entry — profits scale the next bet up, losses scale it down. That is precisely the geometric, fraction-of-wealth structure Kelly's theorem rewards, independent of which fraction is chosen.
- Hard-bounded loss per bet. Each strategy variant runs in its own isolated-margin account: the maximum possible loss of any position, in any scenario, is that account's allocation — never the bankroll. In log-growth terms, this is what keeps ln(wealth) away from minus infinity. Ruin is structurally excluded rather than probabilistically argued away.
- Leverage set by stress-test buffer, not by return optimisation. The leverage tiers in use were chosen by asking "what does the worst recorded dislocation do to this position at this leverage, and how much buffer remains" — a drawdown-constrained sizing rule. That is fractional Kelly in practice: deliberately sizing below the aggressive optimum because the optimum's key parameter is untrustworthy.
- Diversification across parameter tiers, not duplication. Capital is split across variants at different order-book depths — which adds genuine capacity and makes the account-level loss cap meaningful, rather than concentrating the whole bankroll behind one estimate.
The stress-test-derived leverage tier the desk had already settled on months earlier — chosen with no reference to Kelly whatsoever — lands inside the 5–10x plateau that the tail-adjusted log-growth table identifies under the most defensible severity assumption. Two unrelated methods, same answer. That convergence is worth more than either method alone.
5Know when Kelly earns its seat back
"Not usable as a formula today" is not the same as "never usable." The review identified three specific situations where Kelly-style sizing becomes legitimate for this book — each one defined by its assumptions coming back into force:
- Sizing the surge into a known event window. When a dislocation is already underway, the edge is large, short-lived and measurable from past events, and the downside per unit is bounded by the account structure. Deciding how much reserve capital to deploy into that window is a genuine Kelly problem — big measurable edge, bounded loss, definable horizon.
- Carry positions with near-i.i.d. daily outcomes. A hedged funding-rate book produces a long series of small, roughly independent daily accruals with bounded basis risk — much closer to Kelly's native habitat than queue-dependent market-making. Sizing that book as a fraction of bankroll is defensible arithmetic.
- Choosing between discrete leverage tiers at scale-up. Not "solve for optimal leverage" — but "given tiers with known stress buffers, compare their tail-adjusted log growth under a range of severity scenarios," exactly as in the table above. Used as a scenario comparison rather than a point estimate, the log-growth lens is the right tool for that decision.
What didn't happen: no sizing parameter changed because a formula said so. The output of this review is the table in Step 2, the list of broken assumptions in Step 3, and a documented rationale for why the existing structural rules already occupy the defensible region — plus the three defined situations in which the formula graduates from commentary to calculator.
The reusable checklist
Before applying Kelly — or accepting anyone else's Kelly-derived sizing — ask:
- Does the return sample contain the loss events that actually threaten the strategy — or only the calm period between them? An enormous Kelly fraction is a data-completeness alarm, not a recommendation.
- How sensitive is the "optimal" answer to tail assumptions the data cannot pin down? If the answer moves by orders of magnitude across plausible scenarios, the formula is delegating the real decision back to you.
- Are the bets independent draws, or regime-driven? A single global fraction is the wrong shape for a regime-switching book.
- Does bet size alter the odds themselves — through queue position, market impact, or capacity? If so, capacity binds before Kelly does, and the formula's premise is gone.
- Is there a continuous cost of carrying positions (financing, borrow fees) that scales with time rather than bet count? Kelly's discrete-round framing hides it.
- Is ruin structurally excluded (hard loss caps per bet, isolated accounts), or only probabilistically unlikely? Log-growth mathematics treats those two very differently.
- Whatever fraction is chosen — is the structure Kelly-shaped (fixed fraction of live equity, compounding both directions), even if the fraction itself comes from stress tests rather than the formula?
- Do independently derived sizing answers (stress buffers, scenario log-growth, capacity measurements) converge? Agreement between unrelated methods is stronger evidence than precision within one.
Where this fits
This is the same discipline MOA applies to enterprise vendor claims and architecture decisions under Independent Verification & Validation — extended to quantitative sizing and risk frameworks. We review methodology and evidence; we don't sell or manage the strategies under review, and nothing here is personalised financial advice. The value is the same as everywhere else we work: no stake in the answer, no incentive to bless the formula that flatters the plan.