Skip to content
Deuce Ladder

The format

Deuce Rating methodology

This is the full specification of the Deuce Rating: the algorithm, the parameter values, the update rule, a worked example you can check with a calculator, how the confidence figure is computed, and what the rating cannot do.

Methodology version {{METHODOLOGY_VERSION}}, effective {{METHODOLOGY_EFFECTIVE_DATE}}. Reviewed every quarter whether or not anything changed. Every revision is listed in section 12 and cross-posted to the product changelog. A review that finds nothing to change still gets a dated line, because a specification with no review record is a specification nobody can tell the age of.

The rating page explains the idea in a few hundred words. This page is the receipt behind it. It exists because a rating nobody can check is a rating you are being asked to take on faith, and almost nothing in this sport is published in enough detail to make checking possible.

Of eleven tennis rating and league platforms audited in July 2026, one stated in public which algorithm produced its rating. None published a per-player confidence figure.

Source: Deuce Ladder competitive research, audit of eleven tennis rating and league platforms. Captured July 2026, next recheck October 2026. Our own research rather than a third party's, so here is the test applied: a platform counted as naming its algorithm only if its public documentation gave a model by name, in writing, without a login. Proprietary, advanced and results-based did not count. The platform list is available on request.

So if you sit on a club committee, weighing up whether to put your members behind a number they are certain to argue about at some point, this is the page to read and the one to forward around.

One concession before you start, since it is the first thing a careful reader will look for. Deuce has not played a season yet. Every parameter below is an initial choice, made from published work and from the shape of the format, and not a value fitted to real Deuce matches. That and six other limitations are in section 11, and section 11 is the part of this page worth reading first if you are looking for the catch.

One person at Deuce is accountable for this document, and it is not somebody in marketing. Corrections are wanted, not merely tolerated: write to hello@deuceladder.com or use the contact page. Find an arithmetic error in section 4 and it is treated as a defect, with a changelog entry like any other.

{{RATING_ALGORITHM}} IS AN OPEN DECISION. Options, tradeoffs, current lean and the five blocking questions are documented in the long comment on /rating under "The algorithm, by name." That comment is the source of truth for the decision. Do not duplicate it here and do not let the two drift.

THIS PAGE IS WRITTEN SO THE STRUCTURE AND THE COMMITMENTS ARE LOCKED WHILE THE MATH IS NOT.

Locked, regardless of which option wins: - The algorithm is named in public (section 2). - The parameter values are published (section 2). - The update rule is stated in plain language (section 3). - A worked example is hand-checkable with a calculator (section 4). - Confidence is published next to every rating, everywhere (section 5). - Provenance affects confidence, and at launch not the update (section 6). - Decay is continuous with no window boundary (section 7). - Placement reads the top of the plausible range (section 8). - Singles and doubles are separate objects (section 9). - The mapping is approximate and unofficial (section 10). - Limitations are stated plainly (section 11) and every revision is changelogged (section 12).

Tokenized, resolving when the algorithm lands: {{RATING_ALGORITHM}}, {{RATING_SCALE}}, {{RATING_INITIAL}}, {{EXPECTED_SCORE_RULE}}, {{UPDATE_RULE}}, {{UPDATE_PARAMETERS}}, {{PARAMETER_TABLE}}, {{CONFIDENCE_OBJECT}}, {{CONFIDENCE_RULE}}, {{DEVIATION_INITIAL}}, {{DEVIATION_FLOOR}}, {{HALF_LIFE_DAYS}}, {{MARGIN_TERM}}, {{DOUBLES_ATTRIBUTION}}, and every {{WE_*}} value in section 4.

WHEN THE DECISION LANDS: resolve the tokens, fill the worked example with real arithmetic, verify the arithmetic by hand once before publishing, and add a changelog entry. Do not commission a rewrite of this page. If the structure has to change to fit the chosen model, that is a signal the model breaks a published commitment, and the commitment wins.

DO NOT SHIP THIS PAGE WITH ANY TOKEN IN THIS LIST UNRESOLVED. An unresolved token on the page whose entire purpose is verifiability is worse than no page. ============================================================================ -->

1. What the rating is, and what it decides

1.1 The Deuce Rating is a point estimate of playing standard. It is published with a confidence figure that says how much information sits behind the estimate.

1.2 It is expressed on the {{RATING_SCALE}} scale. A new player with no evidence starts at {{RATING_INITIAL}}.

1.3 Your rung is a separate object governed by the rules, applied identically to everyone. The rating sits beside the standings as a measurement and moves no player up or down. The separation is deliberate, and the rating page explains why at greater length.

1.4 Two objects are stored per player per discipline: the rating and {{CONFIDENCE_OBJECT}}. They are always read together, and no code path returns one without the other.

1.5 Every confirmed match under section 9 of the rules updates both objects. A pending result, a disputed result, and a voided match under rule 7.4.3 update neither.

1.6 Walkovers do not update the rating. A walkover moves rungs exactly as a played win does under rule 10.4, and it produces no evidence at all about how anybody plays tennis. Counting it as evidence would pay a player for an opponent's silence.

1.7 Retirements update the rating using the completed portion of the match only, and are marked.

1.8 The rating drives exactly two things inside Deuce, and a club evaluating it should hold us to this list. It places a new player into a level band under section 8. It is offered to an organizer as one input when they seed a band or a playoff bracket, alongside match record and whatever the player brought at signup, and the organizer decides. That is the whole list.

1.9 The rating is not a tiebreaker for rung order and it never will be. Two players who finish a season level are separated by the five-step order published on the scoring page, and every step in it is something both players can recompute from results they can see. A rung that depended on this document would be a rung nobody could check.

1.10 One rating per player per discipline, across everything. A player in three ladders at once has one singles rating, fed by all three, not three ratings that disagree with each other. Rungs are per ladder. The rating is not.

1.11 A result that is reversed is reversed in the rating too. If a confirmed score is later voided under rule 7.4.3, or a falsified score is corrected under rule 17.5, the match is removed from the evidence and every rating downstream of it is recomputed under 3.7. Both players are notified that their rating changed and why, and the correction is visible on the match record. A platform that leaves a falsified result sitting in the arithmetic after removing it from the standings has published a penalty it did not actually apply.

2. The algorithm and its parameters

Two things get published here that competitors keep back: the name of the model, and every constant it runs on. The second one is the part that takes nerve, and it is also the part without which the first one is decoration.

2.1 The Deuce Rating is calculated with {{RATING_ALGORITHM}}.

2.2 The parameter values are published in full. A named algorithm with hidden constants is half a disclosure, and half a disclosure is what this page exists to avoid.

Parameter Value What it controls
{{PARAM_1_NAME}} {{PARAM_1_VALUE}} {{PARAM_1_EFFECT}}
{{PARAM_2_NAME}} {{PARAM_2_VALUE}} {{PARAM_2_EFFECT}}
{{PARAM_3_NAME}} {{PARAM_3_VALUE}} {{PARAM_3_EFFECT}}
Initial rating {{RATING_INITIAL}} Where a player with no evidence starts
Initial uncertainty {{DEVIATION_INITIAL}} How fast a new player's rating moves
Uncertainty floor {{DEVIATION_FLOOR}} The most certain the system will ever claim to be
Result half-life {{HALF_LIFE_DAYS}} days How fast a result loses weight, see section 7
Margin term {{MARGIN_TERM}} Whether and how score margin enters the update

2.3 The parameters are versioned with the methodology. A change to any value in that table is a methodology revision, gets a version number, and appears in section 12.

2.3.1 The table has three named parameter rows because the number of tunable parameters is a property of the model, and the model is the open decision above. Elo has one. Glicko-2 has more. Whatever the count turns out to be, every parameter the model exposes appears in that table with its value. None is held back as commercially sensitive, because a constant you cannot see is a constant you cannot check the worked example against.

2.4 The algorithm gets published because the algorithm was never the moat. The match data is. Publishing the math costs nothing and buys the number the one thing that makes a rating useful, which is your belief in it.

2.5 If the chosen model turns out to be unable to honor a commitment in this document, the commitment wins and the model is the thing that changes. That is written down here rather than left as an internal understanding, because the order of those two priorities is the only real guarantee behind any of the sections below. A model that cannot publish a per-player confidence figure, for instance, is not a model Deuce ships, however well it predicts.

3. The update rule

This is the section a skeptic came for, so it is written to be read twice. The rule appears three times on this page and the repetition is deliberate: in words below, as a numbered procedure at 3.5, then with actual numbers in section 4. Different readers need different versions of the same thing, and none of them should have to guess which one they are looking at.

3.1 In words. Before the match, the model works out how likely each player was to win, using the two ratings and the uncertainty attached to each of them. After the match, each rating moves toward the result by an amount proportional to how surprising the result was, scaled by how uncertain the model was about that particular player. A predictable result barely moves either rating. An upset moves both of them, and it moves the rating of a player the system hardly knows further than anyone else's.

3.2 The expected result. {{EXPECTED_SCORE_RULE}}.

3.3 The update. {{UPDATE_RULE}}.

3.4 The uncertainty update. {{CONFIDENCE_RULE}}. Uncertainty falls when you play, rises when you do not, and falls faster against opponents the system already knows well.

3.5 Order of operations, for one confirmed match:

  1. Read both players' current rating and uncertainty, decayed to the match date under section 7.
  2. Compute the expected result under 3.2.
  3. Apply the margin term {{MARGIN_TERM}}, but only if the match was played in a format on the declared comparable-format list, which is published alongside the parameter table in section 2. A match played in any other format is treated as a binary result, win or loss, with no margin contribution at all.
  4. Update both ratings under 3.3.
  5. Update both uncertainties under 3.4.
  6. Apply the provenance treatment under section 6.
  7. Write the new values, the input values, and the methodology version to the match record.

3.6 Step 7 is what makes this page checkable. Every match record stores the ratings that went in, the ratings that came out, and the version of this document that produced them. A player asking why their rating moved gets the answer off the record instead of out of a support agent's head.

3.6.1 Those stored numbers are yours to read, and asking for them is not a favor. Every player can see the inputs and outputs on their own matches inside the product, and export the whole set under the data-export commitment on the privacy page. An organizer can request the same for every match in their own ladder. Publishing an update rule while keeping its inputs unreadable would make this document checkable in theory only.

3.7 Updates run per match, in the order the matches were played, not the order they were confirmed. A match confirmed late gets inserted at its played date and the chain downstream of it is recomputed. Skip that step and two players end up with different numbers out of the same evidence, decided by which of them happened to press confirm first.

3.8 Recomputation is bounded and the bound is published: a match cannot be inserted more than {{RECOMPUTE_HORIZON_DAYS}} days after the date it was played. Past that point the result still goes on both players' match records, marked late, and it does not enter the rating. The tradeoff is stated so nobody has to guess at it. An unbounded rewrite would let somebody's rating move today because of a match from two seasons ago that neither player remembers clearly, and a horizon costs one data point instead.

4. A worked example

One match, every number shown. Take a calculator to it. The whole point of this section is that you can.

Player A and Player B are placeholders, not people. Deuce is pre-launch and there are no real matches to publish yet. When the first season is played, this example is replaced with a real match record and a link to it.

4.0 Until then, every number below renders from a named test fixture, stamped with the methodology version it was computed under, and the page says so on the fixture's own label instead of leaving you to work it out. The stamp matters: the fixture and the parameter table in section 2 read the same configuration at the same version, and if they ever disagree the page shows both stamps rather than quietly rendering a worked example against constants that have moved. The fixture is executed in the build, so the three tables cannot drift away from the parameter values in section 2, and the arithmetic is reconciled by hand before every publish. The fixture also has to satisfy the description in the paragraph beneath it: Player B has to be the lower rated of the two, because the example's whole job is showing what an upset does. A fixture that violates that fails the build rather than quietly making this page wrong.

Before the match

Player A Player B
Rating {{WE_A_RATING}} {{WE_B_RATING}}
Uncertainty {{WE_A_DEVIATION}} {{WE_B_DEVIATION}}
Confidence shown to players {{WE_A_CONFIDENCE}} {{WE_B_CONFIDENCE}}
Played matches on Deuce {{WE_A_MATCHES}} {{WE_B_MATCHES}}
Days since last played {{WE_A_LAYOFF}} {{WE_B_LAYOFF}}

The match. One set with a tiebreak at 6-6, finishing {{WE_SCORE}}. Player B wins it from the lower rating, which makes this an upset and makes it the useful case to work through. Provenance tier is confirmed, meaning Player B submitted the score and Player A actively confirmed it instead of letting the timer do the job.

The arithmetic

Eight rows, one per step of 3.5. Step 5 is where the uncertainty scaling does its work, and it is the row that explains why an upset against an unknown player moves so much.

Step Rule Input Output
1 Section 7 Ratings decayed to the match date {{WE_A_RATING_DECAYED}}, {{WE_B_RATING_DECAYED}}
2 3.2 Rating difference {{WE_RATING_GAP}} Expected result for A: {{WE_EXPECTED_A}}
3 3.3 Actual result for A: 0. Surprise: {{WE_SURPRISE}} Raw update for A: {{WE_RAW_UPDATE_A}}
4 {{MARGIN_TERM}} Games won and lost: {{WE_GAME_MARGIN}} Margin adjustment: {{WE_MARGIN_ADJUSTMENT}}
5 3.3 Uncertainty scaling for A: {{WE_A_DEVIATION}} Applied update for A: {{WE_APPLIED_UPDATE_A}}
6 3.3 Same, for B Applied update for B: {{WE_APPLIED_UPDATE_B}}
7 3.4 Both players played, one recent result each New uncertainties: {{WE_A_DEVIATION_NEW}}, {{WE_B_DEVIATION_NEW}}
8 Section 6 Provenance tier: confirmed {{WE_PROVENANCE_EFFECT}}

After the match

Player A Player B
Rating {{WE_A_RATING_NEW}} {{WE_B_RATING_NEW}}
Change {{WE_A_DELTA}} {{WE_B_DELTA}}
Uncertainty {{WE_A_DEVIATION_NEW}} {{WE_B_DEVIATION_NEW}}
Confidence shown to players {{WE_A_CONFIDENCE_NEW}} {{WE_B_CONFIDENCE_NEW}}

4.1 Watch what happened to Player B's uncertainty. It fell, because the system knows more about them than it did an hour ago, and that movement is separate from the movement in the rating. The two run on different clocks and it is worth getting used to seeing them disagree.

4.2 Notice also that no rung appears anywhere in these tables. Player B took Player A's rung under rule 10.1 the moment the score was confirmed, and not one step of that depended on the arithmetic above it.

4.3 A second worked example covering a doubles result is published once the attribution question in section 9 is settled.

5. Confidence, and how it is computed

No competitor in this category makes the commitment in this section, so it is the one worth reading slowly. A rating on its own tells you where the system thinks you are. The figure beside it tells you how much of that to believe, and the two are stored, read and displayed as a single object.

5.1 Every Deuce Rating is published with a confidence figure, in the same visual unit, everywhere the rating appears. Profile, standings context, match record, challenge screen, ladder page, API response. There is no surface where a rating appears alone.

5.2 Confidence is derived from {{CONFIDENCE_OBJECT}} and expressed to players in three bands.

Band Threshold Rendering
Low Uncertainty strictly above {{CONFIDENCE_LOW_THRESHOLD}} The rating renders as a range, marked provisional. Never a point estimate with an asterisk
Medium Uncertainty at or below {{CONFIDENCE_LOW_THRESHOLD}} and strictly above {{CONFIDENCE_HIGH_THRESHOLD}} The rating renders with a visible margin
High Uncertainty at or below {{CONFIDENCE_HIGH_THRESHOLD}} The rating renders as a number

5.2.1 Note the direction. Uncertainty falls as confidence rises, so the high-confidence band is the one with the smallest number in it, and a player crossing a threshold downward is moving up a band. Each boundary belongs to the more confident of the two bands it separates, so a player sitting exactly on {{CONFIDENCE_HIGH_THRESHOLD}} is in the high band.

5.3 Three terms feed it: a count of results, a recency weight on each of them under section 7, and the variety term in 5.4. The third one is the term readers usually have not thought about.

5.4 Variety is a computed input and not a figure of speech. Twenty results spread across three opponents contain fewer independent observations than twelve results spread across twelve, and the confidence figure has to reflect that or it is measuring volume instead of information. The variety term is {{VARIETY_TERM}}, computed as {{VARIETY_CALCULATION}}.

5.5 Two matches and forty matches do not produce the same kind of number, and printing both in the same typeface tells a reader they are comparable when they are not. That is the whole argument for this section, and it is the reason 5.1 has no exceptions in it.

5.6 Confidence displays uncertainty and nothing else. It is not a score, it never appears in a leaderboard, and there is no badge attached to it. Nobody is meant to farm the figure, and there is nothing to farm.

6. Provenance weighting

Two identical scorelines can be worth different amounts to a rating. This section covers how much different, and why that difference is displayed at launch instead of being computed into the arithmetic.

6.1 Every confirmed match carries a provenance tier, written at result creation and immutable afterward. Immutable matters here: a tier that could be upgraded later would be a tier somebody could be talked into upgrading. The five tiers and their definitions live on the scoring page, which is the single source of truth for them.

6.2 At launch, provenance moves your confidence figure and leaves the rating update alone.

6.3 The reason for that choice is worth stating. Weighting the update by provenance leaves a rating movement that cannot be explained to the player it happened to. Weighting confidence says the same thing honestly: we know less about a result nobody confirmed.

Tier Effect on the update Effect on confidence
Verified None at launch Full, and the countersignature is shown on the record
Confirmed None at launch Full
Auto-confirmed None at launch Reduced by {{PROV_AUTO_HAIRCUT}}
Contested None at launch Reduced by {{PROV_CONTESTED_HAIRCUT}}, and marked
Unverified Excluded from the update entirely Not counted

6.3.1 The five tiers are an ordered evidence scale, strongest first, and the order on the scoring page is the same order. Verified sits above Confirmed there because a countersigned result is better evidence than a self-reported one that the loser agreed with. At launch the arithmetic does not act on that particular gap: Verified and Confirmed both carry full weight, and the difference between them is displayed on the match record rather than computed into anything. If it ever gains arithmetic effect, that is a methodology revision under section 12, with a version, a date and a statement of what moved.

6.4 Unverified results are imported history and self-reported matches played outside the challenge system under rule 8.8. They sit on your record and they move nothing. A match that was never a challenge is a practice match as far as this document is concerned, however seriously the two of you played it.

6.5 Provenance never affects your rung. A rung that quietly depended on a weighting nobody could see would not be a ladder.

6.6 If this policy changes, both this page and the scoring page change in the same release, with a changelog entry on each.

7. Continuous decay

Every rating system in this sport has to decide what to do about results that are getting old. The usual answer is a twelve-month window, and the usual answer has a hole in it that players find inside a season.

7.1 Every result loses weight continuously from the day it was played. Its weight at age d days is {{DECAY_FUNCTION}}, with a half-life of {{HALF_LIFE_DAYS}} days.

7.2 There is no window boundary anywhere in the implementation. No result "drops out" on a date, and no result counts fully and then counts zero.

7.3 Continuous decay was chosen for a reason and not for elegance. Nearly every rating system in this sport runs a fixed window, and competitive players find the hole in one inside a season. A high rated player whose form has dropped simply waits. The window rolls forward, the good results stay in it a while longer, and they come back carrying a rating they no longer deserve. A fixed window creates a date worth waiting for. Continuous decay never does.

7.4 Uncertainty responds to a layoff faster than the rating does. Time away means the system knows less about you, which is a different statement from saying you got worse.

7.5 So your rating can move slightly on a day you played nothing. That is correct behavior and not a bug. The evidence behind it is a day older than it was.

7.6 Rating decay is not rung decay. Rung decay is a ladder rule with a published rate under section 12 of the rules, and it exists to stop players parking on a good rung. A declared absence pauses that clock. It has no effect on this one, because a notification about your travel plans does not make last February's results any more recent than they are.

7.7 The two systems keep distinct product labels, "rung decay" and "rating confidence", so a support conversation about one is never about the other.

8. Cold start and placement

Everything in this section follows from one rule, which is that being hard to measure must never be an advantage. The rest is arithmetic in service of that.

8.1 A new player with no accepted rating input starts at {{RATING_INITIAL}} with uncertainty {{DEVIATION_INITIAL}}, which is the widest range the system ever publishes.

8.2 A player who brings a rating is seeded from it. Accepted inputs are the canonical list: NTRP, UTR, WTN, LTA rating, ITN, a national federation ranking, or a club level. Every one of them is optional, and a player with none of them answers a few plain questions instead and lands under 8.1.

8.3 An imported rating seeds the rating and leaves the uncertainty alone. Nobody here can verify what you typed in, so an imported number is a starting point and not evidence. {{SEED_TREATMENT}}.

8.4 Placement reads the conservative bound, which means the strong end of the plausible range. Never the middle and never the weak end.

8.4.1 Strong end, not top, because two of the seven accepted inputs run backwards. On NTRP and UTR the stronger player carries the higher number. On WTN and ITN the stronger player carries the lower one, which 10.4 warns about. The rule is about playing standard and not about arithmetic direction, so any implementation that hardcodes a maximum rather than a comparison on standard is wrong even where it happens to produce the right answer. The rating page states the same rule as the top of the plausible range, using an NTRP example where stronger and higher coincide. Same rule, one audience.

8.5 If the system believes you are somewhere between a 3.5 and a 4.5, you are placed as a 4.5 until results say otherwise. In practice that means a new player's first opponent is drawn from the harder end of the guess, and being unmeasured costs you the benefit of the doubt.

8.6 The reason is a design commitment stated publicly on the rating page, and it is stated here as arithmetic. If uncertainty produced easier opponents, the optimal strategy for winning a season would be to remain unmeasured. Post few results, play inconsistently, spread matches across ladders. Every sandbagging problem in amateur tennis is a version of that.

8.7 High uncertainty is what widens the plausible range, and it widens upward faster than downward. An unproven player is assumed capable, never assumed weak. Read together with 8.4, that means the least measured players are placed against the strongest opponents their range allows.

8.8 A player may always ask to be placed higher. Nobody has ever gamed a rating system to get harder matches, so that request needs no defenses.

8.9 The tradeoff is real. A genuinely new player may find their first match or two harder than necessary. Rule 5.4 is the fix: your first challenge in a ladder is unrestricted within your band, once, so you correct a placement downward with one played match instead of waiting for the model to notice.

8.10 This is enforced by test, not by convention. A test fails if band or bracket assignment ever reads the point estimate or the lower bound. That test exists because this is the rule most likely to be softened by a well-meaning ticket about new-player experience.

9. Singles and doubles

Two disciplines and two numbers, with one question still open underneath them: how a result produced by four people gets attributed to each of them. What follows separates the settled part from that.

9.1 Two ratings per player, calculated independently, stored as two objects, displayed separately. There is no blended number and there is no discipline flag on a single object.

9.2 A blended rating ends up describing neither player. Every club has a singles player who is lost at the net, and a doubles specialist whose second serve does not survive a baseline rally.

9.3 Doubles at launch is fixed pairs for the season. The pair is the ranked entity. Rotating partners is on the roadmap and is not built.

9.4 Fixed pairs simplify attribution considerably, because the pair persists across the season and its results accumulate against it. Individual attribution within a pair is {{DOUBLES_ATTRIBUTION}}.

9.5 Your doubles rating does not affect singles placement, and your singles rating does not affect doubles placement.

9.5.1 No per-player doubles rating is displayed at launch. Doubles results are stored and attributed from day one, and the number stays unpublished until 9.4 is settled and written into this section. A rating for the pair may ship before a rating for each partner, since a fixed pair is a stable thing to measure, and this section will say so before the product shows it. Publishing a per-player number this document cannot yet explain would break the commitment the whole page rests on.

9.6 A doubles result records both partners and both opponents. Every one of the four match records stores the inputs and outputs under rule 3.6.

10. Mapping to NTRP, UTR, and WTN

A conversion table is the easiest place on a page like this to overclaim, so most of this section is devoted to what the mapping does not mean. The table itself is at 10.6.

10.1 Deuce publishes an approximate mapping in both directions, as a table of ranges, not a formula.

10.2 What the mapping can claim: that a Deuce Rating in a given range typically corresponds to a given band in another system, across the population we have measured.

10.3 What the mapping cannot claim, and will not be worded to imply:

10.3.1 It is not equivalence. NTRP, UTR, and WTN measure different things, on different scales, with different populations and different verification standards.

10.3.2 It is not official. The USTA, UTR and the ITF have not endorsed it, and Deuce does not use their logos.

10.3.3 It is not reversible. Convert a Deuce Rating to NTRP and back and it will not always land where it started, because the ranges overlap. Anyone claiming a conversion that round-trips perfectly is mapping two scales that were never the same shape.

10.3.4 It is not a substitute. If your league needs an NTRP, you still need an NTRP.

10.3.5 It is not stable yet. Every mapping table is a claim about a population, and this one has no population behind it until real seasons have been played. The first published table is marked provisional and carries the date it was computed.

10.4 Note the direction on WTN. It runs from 40 to 1 and lower is stronger, which confuses everyone exactly once.

10.5 The mapping table is a versioned data component with the computation date rendered beside it. It lives in the dataset and is never hand-maintained in a CMS.

10.6 The table covers all seven accepted inputs from 8.2, not only the three in this section's heading. The heading names NTRP, UTR and WTN because those are the three people search for.

Input Direction Deuce Rating range Confidence in the mapping
NTRP Both {{MAP_NTRP_RANGE}} {{MAP_NTRP_CONFIDENCE}}
UTR Both {{MAP_UTR_RANGE}} {{MAP_UTR_CONFIDENCE}}
WTN Both, inverted scale {{MAP_WTN_RANGE}} {{MAP_WTN_CONFIDENCE}}
LTA rating Inbound only {{MAP_LTA_RANGE}} {{MAP_LTA_CONFIDENCE}}
ITN Inbound only {{MAP_ITN_RANGE}} {{MAP_ITN_CONFIDENCE}}
National federation ranking Inbound only {{MAP_FEDERATION_RANGE}} {{MAP_FEDERATION_CONFIDENCE}}
Club or league level Inbound only, mapped by the organizer {{MAP_CLUB_RANGE}} {{MAP_CLUB_CONFIDENCE}}

10.7 Inbound only means Deuce reads the input and does not publish a conversion running the other way. Deuce will not hand anybody a number and describe it as their LTA rating or their ITN. Those scales belong to the bodies that maintain them, we have measured no population against either one, and a conversion published under those circumstances would be a claim we have no standing to make.

11. Known limitations

Everything above this heading is what the rating does. Everything below it is where the rating stops working, and this is the section to read first if you are looking for the catch.

Seven of them, listed because a methodology page with no weaknesses in it is marketing.

11.1 Ladder volume is low. A club player may post eight matches in a season. Not eighty. Every rating system in tennis is working from thin evidence and ours is no exception, which is why the confidence figure exists. It is an honest response to the problem and not a fix for it.

11.2 Small ladders produce correlated evidence. In a ten-player ladder, everyone plays roughly the same handful of opponents, so the whole ladder's ratings can drift together relative to the outside world. The variety term in 5.4 detects it and reports it as low confidence. It cannot correct it.

11.3 Formats are not comparable. A best of three, an eight-game pro set, and a FAST4 match are not equivalent evidence, and the six accepted formats are all permitted under rule 8.5. The margin term is restricted to comparable formats for this reason, and even then it is a normalization, not a truth.

11.4 Self-officiated results are self-reported results. Nobody at Deuce watched your match. The provenance tiers in section 6 grade how a result was established, which is the most any platform can do without an official on court.

11.5 A rating cannot predict one match. It estimates a standard and states how sure it is about the estimate. On any given Saturday the lower rated player wins often enough, more or less the reason anybody bothers turning up.

11.6 We have no data yet. Deuce is pre-launch. Every parameter in section 2 is an initial choice based on published work and on the structure of the format, not a value fitted to Deuce matches. The first retune after real seasons exist will change numbers on this page, and it will appear in section 12 with the reason.

11.7 A fast improver outruns any decay setting. The confidence figure describes how much information sits behind a rating. It does not describe whether the rating is still true. A junior improving through a season, or an adult coming back from a long injury, can change standard faster than evidence at any half-life can track, and the number will lag reality while reporting perfectly good confidence in itself. Every rating system in club tennis shares this failure. The partial answer on a ladder is that the ladder corrects faster than the rating does: a rung moves the moment a match is confirmed, and rule 5.4 lets a misplaced player reach their real level in one challenge. The rating catches up afterward.

12. Changelog, and what a revision does to your rating

Every change to the rating gets an entry here. Every rating system eventually attracts the same accusation, which is that the math was quietly adjusted to suit somebody, and the only thing that answers it is a dated record that already existed before the question was asked.

12.1 One row per revision, newest first. The fourth column is the one a skeptical organizer reads: it states whether existing ratings were recomputed, and if they were, how many moved and by how much at the median.

12.2 Until there is a first revision the table renders no rows and the notice below it stands in their place. That notice ships as written. It is not the kind of placeholder somebody swaps for a friendlier line, because an empty changelog on a pre-launch product reads as honest and a missing one reads as something being kept back.

Version Date What changed Effect on existing ratings
{{CHANGELOG_VERSION}} {{CHANGELOG_DATE}} {{CHANGELOG_SUMMARY}} {{CHANGELOG_EFFECT}}

No revisions yet. Version {{METHODOLOGY_VERSION}} is the first published methodology. When a parameter changes or history is recomputed, it appears here first and in the product changelog at the same time.

Where to go next

If you came here to check the math. It is all above. If something in section 3 or section 4 does not reconcile, tell us. We will either fix the page or show you the query that produced the number.

If you came here from a club committee. The two commitments that matter for your decision are that the algorithm is named and that every rating is published with its confidence. Both are in the copy above and both are testable.

If you have to circulate this. Print it or send the link. The version number and the effective date sit on the first page of the printed version, which is what a committee paper needs and what a screenshot of a marketing page cannot give you. Nothing on this page is behind a login, so the people you send it to do not need an account to read it or to disagree with it.

If you want the shorter version. The Deuce Rating, explained. How results enter the system. How a ladder works. The vocabulary. Common questions.

If you want to use it. Join a ladder, or start one and see what organizers get.

You pick who you play. We make sure the match actually happens.

The Tennis Rating Algorithm | Deuce Ladder