methodology
Flat Staking vs. Kelly Criterion: Which Actually Wins?
The Kelly criterion is the mathematically “correct” way to size a bet. On our own real, tracked results, it isn’t what we actually use — and the reason why is worth explaining rather than glossing over.
The three methods
Flat staking: every pick gets the same stake, regardless of the edge or odds. Simple, and immune to a common failure mode — sizing errors compound less when the size never changes.
Graduated / tiered staking: bigger perceived edges get a bigger stake, in discrete bands (small / medium / large).
Real Kelly staking (or a fraction of it): stake size is calculated directly from your edge and the odds using the Kelly formula — see our full breakdown of Kelly criterion sizing.
What we actually measured
On our own real staking history, flat 1-unit staking outperformed both graduated staking and real quarter-Kelly sizing — the more aggressively scaled methods produced worse total results despite Kelly being the theoretically growth-optimal approach.
This isn’t a contradiction of the math. Kelly is provably optimal if your probability estimates are as accurate as the formula assumes. It’s also provably punishing when they’re not — Kelly sizing amplifies stakes exactly on the picks where your estimated edge is largest, which is also where estimation error is hardest to catch. If your model is even slightly overconfident on high-edge picks specifically, Kelly sizing turns that overconfidence into oversized losses.
The actual numbers
On our own tracked picks: flat staking finished ahead of graduated (“soft” tiered) staking by roughly 5 units, ahead of an aggressively tiered version by roughly 7.5 units, and real quarter-Kelly sizing came in at −4.4% ROI relative to flat. None of the scaled methods beat simple, boring, equal stakes — despite each one, on paper, being designed to size up on the picks the model liked most.

Want to see why the variance from any single staking method can swing this comparison over a small sample? The Variance & Bankroll Calculator shows the realistic range of outcomes at a given win rate and stake — the same logic applies whether the stake is flat or scaled.
Why this matters more than it sounds like it should
There’s a selection effect worth naming directly: our own picks aren’t a random sample of all possible bets. They exist specifically because a signal — model, analyst, or both — flagged them as having edge. A pick with a claimed 15-percentage-point edge is more likely to be a genuine mispricing than noise, or more likely to be a modeling error that overstates the edge. Flat staking doesn’t care which explanation is true. Kelly sizing bets heavily on the first explanation being correct, every time.
What we actually do
Flat 1 unit per pick, full stop. stake_tier and a quarter-Kelly percentage
are still calculated and shown on the handout — they’re informative about
how strong the model thinks a pick is — but they don’t determine bet size.
The Kelly & Value Calculator still uses the real
formula, because it’s a genuinely useful way to think about a single bet in
isolation; it’s the portfolio behavior across many picks with correlated
estimation error where the theory and our real results diverged.
The honest caveat
This is one measurement, on one pipeline’s picks, over one period. It’s not a universal argument against Kelly staking — professional bettors with longer track records and tighter-calibrated models use fractional Kelly successfully. It’s a specific, real result about our edge estimates’ reliability, and a reason to distrust “the math says X” arguments that skip the step of checking whether the inputs to the math are trustworthy.
Frequently asked questions
- Is Kelly criterion staking bad for sports betting?
- Not inherently — Kelly is mathematically optimal if your probability estimates are as accurate as the formula assumes. The risk is that Kelly sizing amplifies stakes exactly on the picks where your estimated edge is largest, which is also where estimation error is hardest to catch.
- Why did flat staking outperform Kelly on ServedBets' own picks?
- Because the more aggressively scaled methods (graduated and quarter-Kelly) turned any overconfidence in high-edge picks into oversized losses, while flat staking is immune to that specific failure mode by design.
- Does this mean nobody should use Kelly staking?
- No — this is one measurement, on one pipeline's picks, over one period. Bettors with longer track records and tighter-calibrated probability estimates have used fractional Kelly successfully. The result says more about this specific model's edge-estimate reliability than about the formula itself.
- What staking method does ServedBets actually use?
- Flat 1 unit per pick. Stake tier and quarter-Kelly percentage are still calculated and shown for informational purposes, but they don't determine bet size.