Would You Deploy a 3.40 Sharpe Strategy With Only 13 Trades? My 4-Bot NQ Sleeve + the Framework I Used to Decide
Hey all,
Sharing a case study from this week because it forced me to confront the
small-sample problem in a very concrete way — curious how others handle it.
THE SETUP
My combined backtest fleet (319 profitable bots, 2yr 4h bars) produced a
4-bot long Nasdaq futures sleeve this week:
Strategy P&L Sharpe WR MaxDD Trades
NQ_Futures_PutBackratio_CrashHedge $1,860 3.40 69.2% 2.1% 13
MNQ Tech Breakout Reversal $2,862 1.61 65.9% 5.3% 41
Gen2_Nasdaq_AI_Demand_Synthetic $1,464 2.57 61.5% 3.4% 13
NQ26_Tech_Momentum_Accelerator_v2 $450 3.40 69.2% 0.5% 13
Blended: $8,016, Sharpe 2.52 vs 1.49 fleet average.
THE PROBLEM
Three of the four have exactly 13 trades. My own strict filter requires
20 minimum. A 69.2% win rate on n=13 has a 95% CI roughly spanning 46–87%.
Statistically that's a hypothesis, not an edge.
THE FRAMEWORK I ENDED UP WITH
1. Direction filter FIRST — today's institutional signal decides long/short,
then backtest quality ranks what's left (all A+ short bots shelved
regardless of grade)
2. Quality rank second — Sharpe/Sortino/WR/DD as tiebreakers only
3. Size by sample confidence — the 41-trade bot carries the credibility
load; the 13-trade bots carry the convexity load
4. Fractional Kelly only (full Kelly said 52.8% — deployed a fraction)
5. Kill switch armed: 2 consecutive losing months, 3 consecutive losers
on lead, recency flip below 2/3, or BTC breakdown while BTC-NQ corr ≥0.75
THE STRUCTURE (why I like it despite the sample)
The lead is long NQM26 momentum + a put-ratio backspread (sell 1 closer
put, buy 2 further puts). Average loss on scratches: -$1.16. Worst loss:
-$1.30. Largest win: +$6.35. ~5:1 tail ratio. In chop it bleeds a defined
little amount; in a crash the convexity kicks. Fits the current tape where
institutions are long NQ convexity (18-19k call backspreads) while
net-short ES.
THE HONEST CAVEATS
- 2yr 4h bars, approximated fills (not tick-level)
- June was a scratch month across all four (-$107 to -$555)
- Recency is 2/3 everywhere — one more losing month flips all four to
EXCLUDED on my filter
- BTC-NQ correlation at 0.78 means this "equity" sleeve is secretly
long crypto
Questions for the room:
1. Do you hard-fail small samples or deploy-small-and-watch? Where's your
line and why?
2. For options overlays on futures momentum — anyone got live-fill
experience vs backtest fill assumptions on far OTM puts? My
approximation model assumes mid, which I suspect is generous.
Full writeup with the complete stat sheets, selection funnel, correlation
maps and kill-switch dashboard here if useful:
https://www.theorderbookedge.com/p/crash-hedged-momentum-how-a-4-bot
(PSA: the site's pricing goes 5x later today, so if you were ever going
to sub there, today's the last day at current rates. Not my site, just
flagging it.)
Not investment advice — simulated backtests, small samples, live results
will differ (probably downward).

