ServedBets

methodology

Flat Staking vs. Kelly Criterion: Which Actually Wins?

By ServedBets Analytics — the same pipeline that posts every public result on this site, not an outsourced writer.How the model works →
Share

The Kelly criterion is the mathematically “correct” way to size a bet. On our own real, tracked results, it isn’t what we actually use — and the reason why is worth explaining rather than glossing over.

The three methods

Flat staking: every pick gets the same stake, regardless of the edge or odds. Simple, and immune to a common failure mode — sizing errors compound less when the size never changes.

Graduated / tiered staking: bigger perceived edges get a bigger stake, in discrete bands (small / medium / large).

Real Kelly staking (or a fraction of it): stake size is calculated directly from your edge and the odds using the Kelly formula — see our full breakdown of Kelly criterion sizing.

What we actually measured

On our own real staking history, flat 1-unit staking outperformed both graduated staking and real quarter-Kelly sizing — the more aggressively scaled methods produced worse total results despite Kelly being the theoretically growth-optimal approach.

This isn’t a contradiction of the math. Kelly is provably optimal if your probability estimates are as accurate as the formula assumes. It’s also provably punishing when they’re not — Kelly sizing amplifies stakes exactly on the picks where your estimated edge is largest, which is also where estimation error is hardest to catch. If your model is even slightly overconfident on high-edge picks specifically, Kelly sizing turns that overconfidence into oversized losses.

The actual numbers

On our own tracked picks: flat staking finished ahead of graduated (“soft” tiered) staking by roughly 5 units, ahead of an aggressively tiered version by roughly 7.5 units, and real quarter-Kelly sizing came in at −4.4% ROI relative to flat. None of the scaled methods beat simple, boring, equal stakes — despite each one, on paper, being designed to size up on the picks the model liked most.

Bar chart comparing staking methods: flat staking as baseline, soft tiered staking at minus 5.24 units, aggressive tiered staking at minus 7.52 units

Want to see why the variance from any single staking method can swing this comparison over a small sample? The Variance & Bankroll Calculator shows the realistic range of outcomes at a given win rate and stake — the same logic applies whether the stake is flat or scaled.

Why this matters more than it sounds like it should

There’s a selection effect worth naming directly: our own picks aren’t a random sample of all possible bets. They exist specifically because a signal — model, analyst, or both — flagged them as having edge. A pick with a claimed 15-percentage-point edge is more likely to be a genuine mispricing than noise, or more likely to be a modeling error that overstates the edge. Flat staking doesn’t care which explanation is true. Kelly sizing bets heavily on the first explanation being correct, every time.

What we actually do

Flat 1 unit per pick, full stop. stake_tier and a quarter-Kelly percentage are still calculated and shown on the handout — they’re informative about how strong the model thinks a pick is — but they don’t determine bet size. The Kelly & Value Calculator still uses the real formula, because it’s a genuinely useful way to think about a single bet in isolation; it’s the portfolio behavior across many picks with correlated estimation error where the theory and our real results diverged.

The honest caveat

This is one measurement, on one pipeline’s picks, over one period. It’s not a universal argument against Kelly staking — professional bettors with longer track records and tighter-calibrated models use fractional Kelly successfully. It’s a specific, real result about our edge estimates’ reliability, and a reason to distrust “the math says X” arguments that skip the step of checking whether the inputs to the math are trustworthy.

Frequently asked questions

Is Kelly criterion staking bad for sports betting?
Not inherently — Kelly is mathematically optimal if your probability estimates are as accurate as the formula assumes. The risk is that Kelly sizing amplifies stakes exactly on the picks where your estimated edge is largest, which is also where estimation error is hardest to catch.
Why did flat staking outperform Kelly on ServedBets' own picks?
Because the more aggressively scaled methods (graduated and quarter-Kelly) turned any overconfidence in high-edge picks into oversized losses, while flat staking is immune to that specific failure mode by design.
Does this mean nobody should use Kelly staking?
No — this is one measurement, on one pipeline's picks, over one period. Bettors with longer track records and tighter-calibrated probability estimates have used fractional Kelly successfully. The result says more about this specific model's edge-estimate reliability than about the formula itself.
What staking method does ServedBets actually use?
Flat 1 unit per pick. Stake tier and quarter-Kelly percentage are still calculated and shown for informational purposes, but they don't determine bet size.

See today's picks live

Every model output, tracked and discussed in real time — join the 3,000+ member community where it all happens.

Join Discord