Skip to main content

Trust Center

How every figure on this site is computed

Every number Reach120 publishes about itself is listed here with the rule that produces it and the date it was measured. Where a sample is too small to support a claim, this page says so and shows no number — which is the state most of it is in today.

The four rules every figure follows

These are mechanical, not editorial: each one is enforced by a test that fails the build rather than by an intention to be careful.

  • The noun is checked, not just the numberEach figure records exactly what one unit of it counts, and the wording on every surface may describe that and nothing else. A person who read a blog post is a reader; calling them a student is the same defect as inflating the count, only harder to notice.
  • Nothing is typed by handPublished figures resolve from the database when a page is served. A number typed into a component is true on the day it is typed and wrong forever afterwards, with no mechanism able to notice — which is exactly what happened here before this page existed.
  • A failed measurement shows nothingIf a figure cannot be read while the page is being served, it is absent. It is never defaulted to zero, never carried over from an earlier request, and never estimated.
  • A figure that cannot be recomputed expiresOne figure comes from Google Search Console, which this site has no live credential for, so it is recorded as a dated snapshot with the window it covers. It stops being published automatically once it outlives that window, and the build fails rather than serving a stale one.

The figures published about Reach120

Each one below is a complete count of the rows that qualify, with company accounts and automated test accounts excluded.

  • 559learners have signed upRecomputed live
    One unit of this figure is
    learner accounts registered on Reach120, excluding admin and company-domain accounts.
    How it is computed
    Production database, public_measured_stats() → learners. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    August 5, 2026 (UTC)
  • 223of them joined in the last 90 daysRecomputed live
    One unit of this figure is
    learner accounts registered in the last 90 days, same exclusions.
    How it is computed
    Production database, public_measured_stats() → learners_joined_90d. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    August 5, 2026 (UTC)
  • 4,241Writing and Speaking responses scoredRecomputed live
    One unit of this figure is
    Writing and Speaking responses that have been scored by AI, excluding internal accounts.
    How it is computed
    Production database, public_measured_stats() → ai_scored_responses. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    August 5, 2026 (UTC)
  • 123countries it has been found from in Google SearchDated snapshot
    One unit of this figure is
    countries from which at least one person clicked through to the site from Google Search.
    How it is computed
    Google Search Console, property sc-domain:writing30.com (the domain that still carries the organic traffic for this product), countries with >= 1 click. A count over a fixed window, recorded with its measurement date. It stops being published once it outlives that window.
    Measured
    August 4, 2026 (UTC) · covering the 90 days to 2026-08-04

Outcome evidence: real test results learners report

Learners whose booked test date has passed can tell us what they actually scored, and separately choose whether that result may be counted in a published figure. Nobody sees an official score report: every result here is a number a learner typed, and it stays labelled as reported.

0reported results that may be counted in a published figure — consented, not withdrawn
0reported results stored in total, including those kept private to the learner
0learners who have been asked, after a booked test date passed
0of them answered, counting “not now” and “never ask again”

Not enough data yet — nothing is published from this sample

No average, agreement figure or outcome claim is computed until at least 30 consented results have been reported. None have been reported so far. That floor is a convention rather than a guarantee: below a few dozen self-selected reports, an average describes who chose to answer more than it describes how anyone performed, and no arithmetic repairs that.

What this sample is not yet enough for

  • How closely a Reach120 practice band matched the official bandConsented reported results paired with the practice responses scored before the same sitting, broken down by the model that scored them and reported with the sample size behind each cell.
  • Where the practice estimate runs high and where it runs lowThe same pairs, split by band, so the failure modes are published alongside the agreement rather than averaged into it.

Counted August 5, 2026 (UTC). Accounts belonging to the company and to its automated test runs are excluded from every number above.

How accurate the scoring is, per model

Every score produced since 2026-08-04 records the model that produced it. Paired with the results learners report afterwards, that is how closely a practice band matched an official one — per model, with the sample size behind each figure and the bias the scorer is already known to have.

Not enough data yet — no accuracy figure is published for any model

A pair needs two things at once: a practice response scored by a named model, and a reported official result from that same learner, consented to being counted. There are 37 responses that name the model that scored them and 0 consented results, which pair into nothing yet. Until they do, no agreement figure exists to publish and none is estimated.

Responses that name the model that scored them

  • gpt-4o-mini-2024-07-18 · Writing: 24 responses from 3 learners.
  • gpt-4o · Speaking: 6 responses from 2 learners.
  • claude-cli · Writing: 3 responses from 1 learner.
  • claude-sonnet-4-6 · Speaking: 3 responses from 3 learners.
  • gpt-4o-mini · Speaking: 1 response from 1 learner.

A further 4,259 scored responses record no model, because the scorer was only stamped onto new responses from 2026-08-04 and older rows were deliberately not backfilled with a guess. Those cannot be attributed to any model and are excluded from everything above. Reading and Listening never appear here: both are graded against an answer key, so no model produces their bands. Computed August 5, 2026 (UTC).

Known bias in the practice estimate you get today (gpt-4o-mini)

Two attempts were made to correct the behaviour below and both were measured, found not to work, and reverted rather than shipped. So this is how the scorer behaves in the product right now, and publishing it is not an admission — it is the point.

The estimate pulls toward the middle of the scale
Weak responses are marked up and strong responses are marked down at the same time, so the spread of estimates is narrower than the spread of real ability. Every one of the held-out responses was scored between 1.0 and 4.5, and most between 2.5 and 4.5.
Low bands are marked about half a band to a full band high
On the held-out set, responses written to band 2 were estimated about 0.8 of a band high and responses written to band 3 about 0.6 high. If you are early in your preparation, the practice estimate is more likely to flatter you than to discourage you — which is the direction that matters most, because it is the one that could send someone into a test they are not ready for.
The top band is marked about a band low, and is effectively unreachable
Responses written to band 5 were estimated about 0.9 of a band low, and the live prompt awarded a 5.0 in none of eight attempts at it. A strong writer should read a 4.0 or 4.5 here as compatible with a higher official band, not as a ceiling.
Typical distance from the intended band
Across the held-out set the estimate sat about 0.59 of a band away from the band the response was written to, in either direction — mean absolute error, measured twice with the same result.
What this was measured against
These responses were written and banded to the published ETS descriptors by this project, not scored by ETS. So this measures how consistently the scorer applies our reading of the rubric. It is not a measurement of agreement with an official result, and nothing here should be read as one.

Measured August 4, 2026 (UTC). Source: docs/reach120/32-SCORING-MODEL-BENCHMARK.md §11 — 40 held-out responses written to the ETS band descriptors, scored twice by the live production prompt.

What is deliberately not computed

A figure that is not published is not measured either, so it cannot quietly start being published.

  • How many people have paidNot published and not counted. It depends entirely on what counts as having paid, and a figure whose definition moves is not evidence of anything.
  • Speaking as its own figureScored Speaking responses are counted inside the scored-response total rather than advertised separately.
  • Anything about improvementNo figure here says practice on Reach120 raised anyone’s score. Establishing that needs the outcome sample above, paired with practice history, and it does not exist yet.

If a figure here looks wrong

Tell us and it gets checked against the query that produced it.

Every claim on this site is registered with its evidence and its measurement date, and a claim that fails re-verification is corrected or withdrawn rather than left standing. Contact us with the figure and where you saw it. The Trust Center covers how scoring works and what Reach120 does not claim.