tennis
Why Tennis Ratings Need to Be Surface-Specific
Two players can have nearly identical overall records and be completely different propositions depending on the surface. A single rating number hides that. A surface-specific one doesn’t.
What Elo is, briefly
Elo is a rating system originally built for chess: every result updates both players’ ratings, with upsets moving the number more than expected results do. Applied to tennis, a higher-rated player is expected to win more often, and the gap between two ratings converts directly into a win probability.
The part that matters for tennis specifically: Elo assumes rating is a single, transferable skill level. That assumption breaks the moment you introduce surfaces.
Why one number isn’t enough
A heavy topspin baseliner can be genuinely elite on clay — where the ball sits up and rewards patience — and merely solid on grass, where low, fast bounces reward flatter, more aggressive strikers. A serve-and-volley-adjacent player can show the reverse pattern. Averaging these into one “overall” rating doesn’t produce a meaningful middle ground — it produces a number that’s wrong for every individual surface.
We compute three separate Elo tracks per player — hard, clay, and grass — alongside an overall figure, using a 90-day rolling window rather than a career-long average, so it reflects current form rather than a player’s peak from three years ago. You can see the current numbers on the Elo Rankings page, filterable by surface.
How a rating actually updates
After each match, both players’ ratings move — the winner’s up, the loser’s down — by an amount that depends on how surprising the result was. Beating a much higher-rated opponent moves your rating a lot; beating someone you were expected to beat moves it only slightly. This is what lets Elo track current form faster than a simple win-percentage would: a big upset registers immediately, instead of getting diluted across a season’s worth of matches.
Our model retrains on a rolling 90-day window rather than updating incrementally forever, currently built from roughly 16,000 recent matches across tours. That’s a deliberate tradeoff: it forgets old form faster than a career-long Elo would, which is the right call for match analysis (you care about a player’s level now, not three years ago) but means the exact numbers can shift meaningfully from one training run to the next as the window rolls forward — expect them to move, that’s the system working as intended, not noise.
Reading a rating gap as a probability
The whole point of an Elo-style system is that the gap between two ratings converts directly into a win probability — that’s what makes it useful for value betting rather than just a ranking curiosity. A gap of roughly 100 points corresponds to a moderate favorite (somewhere around 60-65% depending on the exact scale); a gap of 300+ points is a heavy favorite.

The Elo Rankings page shows the raw numbers — the methodology page covers how those numbers feed into an actual fair-probability estimate for a specific match.
What this changes in practice
Take a hard-court match between two players with similar overall ratings. If one of them is a clay specialist whose hard-court Elo is meaningfully lower than their overall number, the overall rating alone would overstate their chances on this particular surface. The surface-specific rating corrects for exactly that gap.
This is also why head-to-head record needs the same caveat: two players’ past meetings on clay tell you very little about how they’ll match up on grass, even if the H2H record itself looks lopsided.
The honest limitation
A 90-day window reacts fast to current form, but it also means a player coming back from a long injury layoff, or one who’s only played a handful of matches on a given surface this window, will have a rating built on a thinner sample — the leaderboard requires a minimum of 15 tracked matches to appear for exactly this reason. Thin samples produce numbers that look precise and aren’t.
See the live ratings on the Elo Rankings page, or read about how this feeds into value calculations.
Frequently asked questions
- What's the difference between overall Elo and surface-specific Elo in tennis?
- Overall Elo averages a player's level across all surfaces into one number, which can misrepresent their chances on any single surface. Surface-specific Elo tracks hard, clay, and grass separately, since a player's level can differ substantially by surface.
- How often is the Elo rating updated?
- The model retrains on a rolling 90-day window rather than a career-long average, so it reflects current form. That also means the exact numbers can shift meaningfully between training runs as the window rolls forward — that's expected behavior, not noise.
- How many matches does the Elo model use?
- Currently built from roughly 16,000 recent matches across tours, with a minimum of 15 tracked matches required per surface before a player appears on the leaderboard, to avoid thin-sample ratings that look precise but aren't.
- How much does a 100-point Elo gap actually mean?
- Roughly a 60-65% win probability for the higher-rated player, depending on the exact scale — see the Elo Rankings page for the live numbers and the gap-to-probability chart in this article.