Every training app claims it works. We ran the numbers on our own users, spent a day trying to kill the result five different ways, published it — and then, two re-runs later, caught ourselves computing the headline wrong. Here's the corrected number, the mistake that produced the old one, and everything that still stands.
We retired our own headline stat. The first two versions of this article claimed active users improved ≈1.7× faster per game than their own pace before joining. That figure was median(post-signup pace) ÷ median(pre-signup pace) — a ratio between two separately sorted lists, not a comparison of each person against themselves. Computed correctly, one difference per user, the lift is about zero: −0.21 points per game at the July run, −0.20 in August, +0.004 now, with only 46% / 49% / 50% of users beating their own prior pace. It was negative on the day we first published it.
We've replaced it with a comparison that holds the number of games fixed, and it gives a real but far smaller edge: +7 median rating points per user, 56% beating their own prior run. The absolute gains were never affected — they were computed per user from the start. The full correction is here.
Interactive lesson — The hanging piece: the question that saves games: Play it on the board below, and make the key move yourself. White to move. Run the cheapest check in chess before you touch anything, which black pieces have a defender, and which have none? Rxd5. Nothing guards the knight, so the rook simply takes it. No calculation, just counting. Now look for the recapture. The pawn on b6 covers a5 and c5, not d5, and the king is on the other side of the board. A whole piece, won by asking one question. Ask it every move, and ask it for your own pieces too.
Third run, Sep 7, 2026: every user with ≥20 post-signup games in their dominant time class, n=167 · deltas computed within that class only · internal accounts excluded. Runs 1 and 2 (n=28, n=59) are in the run table.
Rating change from each user's first to last post-signup game, inside their dominant time class. Every user in the sample is on this chart — including all 42 who lost rating.
This is the whole sample, not a highlight reel. 122 of 167 users gained rating; the spread runs from +552 down to −254, and the middle of the distribution — the number a skeptic should care about — is +35. The biggest gainer (1073 → 1625 in five days) is an account climbing fast on a small number of games, exactly the profile you'd flag as a still-calibrating rating. We flag it too: it's one reason the headline here is a median with a sample size attached, not the mean (+48) it inflates.
That blended median also understates what a settled user sees, because most of this sample is new. Among the 50 users we have now been measuring for 18 days or more, the median is +55 and 84% gained; among the 117 measured for less, it's +26 and 68%. The full split is in the run table, and the reason it matters is arithmetic, not spin: absolute rating gain scales with how long you've been playing.
For context on what those users are actually fixing while they climb, the companion reports cover what each Elo level struggles with and how chess games are actually lost — short version: games are decided by one collapse, and shrinking that collapse is where the rating points live.
Published Jul 13 and Aug 3, 2026. Retracted Sep 7, 2026. Here is the error in full, because a page arguing about honest denominators does not get to quietly edit its own numbers.
The first two versions of this article carried a second headline next to the rating gain: active users improved ≈1.7× faster per game than their own pace before joining. It was the number we told people to watch. It was wrong. Not the data, the arithmetic on top of it.
We computed it as median(post-signup pace) ÷ median(pre-signup pace). Those are two separately sorted lists. Dividing one by the other tells you how two groups compare; it says nothing about what happened to any person, which is exactly what "faster than their own past" claims. The correct computation is one difference per user, then the median of those differences. Do that, and the effect disappears.
| Run | Users | Ratio of medians (what we published) | Paired median lift (correct) | % beating their own prior pace |
|---|---|---|---|---|
| Jul 13, 2026 | 28 | 1.61× | −0.21 pts/game | 46% |
| Aug 3, 2026 | 57 | 1.41× | −0.20 pts/game | 49% |
| Sep 7, 2026 | 135 | 1.49× | +0.004 pts/game | 50% |
Read the last two columns. The paired lift was negative on the day we first published the claim, and it is indistinguishable from zero now, with almost exactly half the sample beating their own prior pace — a coin flip. The two methods diverge because of regression to the mean: users who arrive on a hot streak cool off, users who arrive on a cold one heat up, so the two sorted lists keep their gap while the per-person differences collapse onto zero. The July and August rows are genuine as-of replays — we capped both game dates and signup dates at each run date — so those are the cohorts we published, recomputed, not a convenient new sample.
Nothing else in the article moved. The absolute gains (+35 median, 73% gaining, the mature and fresh split) were computed one user at a time from the start and are unaffected. What we lost is the per-game multiple, and we've replaced it with a comparison that can't make the same mistake — directly below.
Median rating change over a fixed number of games: each user's post-signup run against their own equally long stretch immediately before it. Same person, same game count, same time class (n=135).
The obvious objection to any before/after stat is that maybe these players were improving anyway. So we compare them to themselves — carefully, because the careless version of this is exactly what we got wrong. Take each user's post-signup games in their dominant time class, count them, then take that same number of their games immediately before signing up. One difference per user, no sorted lists.
The median user gained +39 across their post-signup run, against +14 across their own previous run of identical length. The gap between those two medians (+25) is still not the effect — pair it person by person and the median lift is +7 rating points, with 56% of users beating their own prior stretch against a 50% coin-flip baseline. That is a real edge and a small one, and it is roughly an order of magnitude less impressive than the number we used to print.
| Run | Users | Their previous N games | Their first N games after joining | Paired median lift | % who beat their own prior run |
|---|---|---|---|---|---|
| Jul 13, 2026 | 28 | +22 | +57 | +4.5 | 50.0% |
| Aug 3, 2026 | 57 | +15 | +30 | +15.0 | 52.6% |
| Sep 7, 2026 | 135 | +14 | +39 | +7.0 | 56.3% |
The encouraging thing in that table is the direction of the last column: 50.0% → 52.6% → 56.3% across three runs. Three points is not a trend yet, it's a thing to keep measuring, which is why the table exists rather than a single number.
And here's the claim we killed first, back in July: measured per month, the same data says users improve 3.6× faster after joining. That number is real, spectacular, and dishonest. Users simply play more after joining an improvement app, so a per-month framing smuggles play frequency in as skill. Fixing the denominator was right; we then got the arithmetic on the fixed version wrong too, which is a decent argument for re-running your own numbers on a schedule. If you're comparing tools, track your progress in per-game terms — and pair it against yourself, not against a median.
Per-game rating pace — our users before they signed up, versus what published data implies is typical at each level. The grey bars are estimates (caveats below); our number is measured directly.
Comparing someone to their own past only means something if that past was normal to begin with. It was. Our users' pre-signup pace of 0.40 rating points per game lands right on what the wider data implies for their level. Published medians from Chess.com's SmarterChess analysis put a studying 800-rated player near +239 rating a year, and Lichess's public rating data (hundreds of millions of games) shows beginners around 800–1000 gaining their first ~100 points within a few months. Convert either into a per-game figure for an active player and you land in the 0.3–0.5 range — exactly where our users started before joining.
Two honest caveats, because this is a comparison. First, improvement gets harder as you climb: the same external data has a 1600-rated player gaining only about +35 a year, so the grey bars drop fast with level — and our sample skews toward beginners, whose ceiling for quick gains is naturally higher. Second, turning an annual figure into a per-game one means assuming a game volume — the exact move we refused to make with our own per-month stat. So treat the grey bars as a rough band, not a precise line. What is not an estimate is the 0.40: it's counted per game, directly, on real accounts. Earlier versions of this chart carried a second teal bar for our post-signup pace, with the gap between the two as the punchline. We removed it in September 2026 — putting two medians of two different distributions side by side is precisely the error this article now documents, and a chart is no place to repeat it.
Before publishing in July, we attacked the result the way a skeptical reviewer would. The attacks and verdicts, unedited. A sixth attack, run in September, killed a claim these five let through — that one gets its own section.
Blitz and rapid ratings live on different scales; compute a delta across them and you fabricate numbers. True — and our very first pass did exactly that, producing a median of +36. Re-run within each user's dominant time class only, run 1's median rose to +50. The attack was correct, fixing it made the stat stronger, and every run since has used the class-clean rule.
Partially lands. Per month, post-signup improvement looks 3.6× faster — but decomposing it shows most of that is play frequency, not play quality. We demoted the claim to a per-game version, ≈1.7×, and retired the per-month framing permanently. Update, Sep 2026: the per-game replacement was itself wrong, a ratio of medians rather than a paired comparison, and we retracted it too. What survives is an equal-length paired comparison worth about +7 rating points per user.
True in both directions, so we keep them. A +552 climber stays in run 3's sample and so does a −254 collapse; the only account ever excluded was a −661 fall over five days of bullet in run 1, the pattern of account sharing rather than chess. Verdict: the mean (+48) is decorative. The median (+35) is the stat, and the full distribution is charted rather than summarised.
Correct, and undismissable. Requiring 20+ post-signup games self-selects motivated players; users who churned never enter the sample. That's survivorship bias, and without a randomized control group — which we don't have — it can't be argued away. So it goes in the article, not in a footnote: this is a stat about engaged users, not about everyone who signs up.
We haven't separated this effect, and our biggest gainer (352 → 701) pattern-matches a new account still finding its level — see regression toward the mean. It's unlikely to explain the whole distribution, but it's one more reason the headline is a median with n attached, not a hero number.
We will claim: among active users, 73% gained rating within their first weeks (median +35, n=167; +55 and 84% among the 50 measured for 18 days or longer), and that over the same number of games the median user did +7 rating points better than their own immediately preceding run, with 56% beating their own prior stretch. Both are descriptive, both are computed one user at a time, and both come with their sample size attached.
We won't claim: “Chess DNA adds X rating” — that's causal, and observational data can't support causality no matter how much a marketing page wants it to. We won't claim “3.6× faster improvement,” because we know exactly which confound produces it. And we no longer claim “≈1.7× faster per game,” because we checked our own arithmetic and it didn't hold. When any training tool shows you an improvement stat, ask which attacks they ran against it, what denominator the claim is hiding, and whether the comparison was made person by person or list against list. If you'd rather test it on your own games, Chess DNA analyzes them free and shows you the same per-game curve we used here.
A static stat is easy to cherry-pick; a re-run schedule is not. We re-run the identical query monthly — same ≥20-game bar, same dominant-time-class rule, same exclusions. Every run gets a row here, including the ones that dip and the one that caught us out.
| Run | Users | Median gain | % who gained | Median window | Per-game claim, as published |
|---|---|---|---|---|---|
| Jul 13, 2026 | 28 | +50 | 82% | 18 days | ≈1.7× per game |
| Aug 3, 2026 — everyone | 59 | +30 | 75% | 10.5 days | ≈1.6× per game |
| Aug 3, 2026 — mature window (≥18 days) | 17 | +96 | 88% | ≥18 days | — |
| Sep 7, 2026 — everyone | 167 | +35 | 73% | 9.6 days | +7 pts, paired |
| Sep 7, 2026 — mature window (≥18 days) | 50 | +55 | 84% | 32.8 days | — |
| Sep 7, 2026 — fresh (<18 days) | 117 | +26 | 68% | 5.7 days | — |
Read the blended rows honestly and every new run looks like a dip: +50, then +30, then +35. The reason is boring, and it's not the app: each run pulls in a wave of fresh signups whose measurement windows are short — the median is down to 9.6 days now — and absolute rating gain scales with time on the clock. Split by window and the story is stable. Users measured over the same horizon as run 1 (≥18 days) show +55 with 84% gaining in run 3, against +96 and 88% in run 2 on just 17 users, and +50 and 82% in run 1. The fresh two-thirds sit at +26, which is what "hasn't had time yet" looks like in data.
And the stat this article originally told you to watch — per-game pace against your own pre-signup self — didn't move between runs either. It just wasn't measuring what we said it was. Run 3 is where we checked it properly and retracted it. The replacement, a paired equal-length comparison, has been backfilled onto all three runs so the series stays comparable: +4.5, +15.0, +7.0 rating points of median lift, and 50.0% → 52.6% → 56.3% of users beating their own prior run. No new accounts needed excluding this run, and the five attacks in the grill apply unchanged, including the selection-bias disclosure. That's the whole point of re-running: the flattering number moved with cohort mix, and the honest number turned out to need honest arithmetic too.
Users who gained rating — the full sample is charted above, including all 42 who lost.
Median measurement window. Among users measured 18 days or longer, the median gain is +55.
Median rating points gained versus each user's own equally long stretch before joining (n=135).
Share who beat their own prior run of the same length. A coin flip would be 50%.
For engaged users, measurably and modestly, yes. Across 167 Chess DNA users with 20+ post-signup games, 73% gained rating (median +35); among the 50 measured for 18 days or longer, 84% gained (median +55). The honest size of the edge is small: over the same number of games, the median user did +7 rating points better than their own immediately preceding run. No app can claim causality from observational data — motivated players self-select into every such sample.
Median +35, first to last post-signup game, within each user's dominant time class. 122 of 167 gained; the spread ran +552 to −254 over a median 9.6-day window. Because most of the sample is new, that blended figure understates a settled user: for the 50 users measured 18 days or longer, the median is +55 with 84% gaining. The mean (+48) is inflated by outliers, so we lead with the median.
Partly, and we say so in the body: requiring 20+ games self-selects motivated users, and churned users never qualify. That's survivorship bias, and it's undismissable without a control group. Everything we could test (class mixing, play frequency, outliers), we tested. And in September 2026 we caught a genuine error in our own headline stat and published the correction rather than quietly editing the number.
Per-month framing smuggles in play frequency: users play more after joining an app, so monthly gains jump even if each game teaches nothing. Our per-month number (3.6× faster) looked spectacular and we refused to publish it. Per game is the honest denominator — but the denominator alone isn't enough. Our first per-game statistic divided one median by another, which compares groups rather than people. The version that survives pairs each user against their own equally long prior run.
Every user with 20+ games after signup in their dominant time class (n=167, Sep 7, 2026), internal accounts excluded. Delta = last minus first post-signup rating within that class only, since mixing blitz and rapid ratings fabricates deltas. The paired comparison takes each user's post-signup run and their own equally long stretch immediately before it (n=135), giving one difference per user rather than a ratio between two medians.
We re-run the identical query monthly and publish every run. The sample has gone 28 → 59 → 167 users. The blended median moved +50 → +30 → +35, driven by cohort mix rather than results: each run adds fresh signups with short measurement windows. Hold the window fixed at 18 days or more and it's stable at +50, +96 and +55, with 82%, 88% and 84% gaining. Run 3 also caught a methodology error in our own per-game claim, which we retracted in full.
Want the same measurement on your own chess? Chess DNA analyzes your games, finds the mistakes that actually cost you rating, and tracks your per-game pace — the number this article just argued is the only one worth watching.
Method. Sample: every Chess DNA user with ≥20 games played after signup (proxied by their first imported game) within their dominant time class, n=167 as of Sep 7, 2026. Internal accounts excluded. Ratings come from the user's own game records; deltas are last − first post-signup rating computed within that dominant class only — mixing blitz and rapid ratings fabricates deltas (run 1's mixed-class version gave +36 against a class-clean +50). The paired comparison uses the 135 users who also had at least as many pre-signup games in that class: each user's post-signup run against their own equally long immediately preceding stretch, one difference per user. Historical runs are reproduced as as-of snapshots, capping both game dates and signup dates at the run date. The ≈1.7× per-game figure published in July and August 2026 was computed as a ratio of two medians and has been retracted; see the correction. No causal claim is made; see the grill section for the attacks we ran. Part of a series with How Chess Games Are Actually Lost and What Each Elo Level Struggles With.