Skip to main content
20% off your first payment$5.99 for your first week, then $7.49 · Ends November 15, 2026See plans

Trust Center

How every figure on this site is computed

Every number Reach120 publishes about itself is listed here with the rule that produces it and the date it was measured. Where a sample is too small to support a claim, this page says so and shows no number — which is the state most of it is in today.

Who writes this

Every practice page and tool on Reach120 bylines the 'Reach120 research team' and links back here — this is what that means today.

No named, credentialed reviewer has been assigned to Reach120's public content yet, so the honest byline is the team producing it, not an invented person with an invented title. What that team commits to is mechanical, not a credential: every published claim is tracked against its own evidence and review date internally, every ETS attribution carries a source URL and a read date on the page itself, and every quantity on this page and every page it links to resolves through the rules below — never a typed number. When a named, credentialed reviewer is assigned, this section is replaced with their name and a real bio, not added alongside the team byline.

How this content is written

Reach120's practice pages, guides and blog posts are drafted by AI agents. This is what that does and does not mean.

The drafting is automated. Software agents write against a written specification, and the output has to pass mechanical gates before it can ship: every quantity must resolve through the measurement rules below rather than be typed by hand, every ETS attribution must carry a source URL and the date it was read, the scoring and format rules are pinned to the published 2026 specification, and a page that fails any of these fails the build. Structure is gated too — a post that ships as an undifferentiated wall of text is rejected before publication rather than corrected afterwards.

What automation is not allowed to do here is invent. No agent may create a person, a quotation, a testimonial, a rating or a review, and none may present a Reach120 estimate as something ETS published. Where the honest answer is that a number is unknown or a sample is too small, the page is required to say that and show nothing — which is why several figures on this page are absent rather than estimated.

The reason the content exists is to be useful to somebody preparing for the 2026 TOEFL, and the test we hold it to is whether a reader who arrives directly — with no search engine involved — gets what they came for. Pages are written to be read, then found; not the other way round.

The four rules every figure follows

These are mechanical, not editorial: each one is enforced by a test that fails the build rather than by an intention to be careful.

  • The noun is checked, not just the numberEach figure records exactly what one unit of it counts, and the wording on every surface may describe that and nothing else. A person who read a blog post is a reader; calling them a student is the same defect as inflating the count, only harder to notice.
  • Nothing is typed by handPublished figures resolve from the database when a page is served. A number typed into a component is true on the day it is typed and wrong forever afterwards, with no mechanism able to notice — which is exactly what happened here before this page existed.
  • A failed measurement shows nothingIf a figure cannot be read while the page is being served, it is absent. It is never defaulted to zero, never carried over from an earlier request, and never estimated.
  • A figure that cannot be recomputed expiresOne figure comes from Google Search Console, which this site has no live credential for, so it is recorded as a dated snapshot with the window it covers. It stops being published automatically once it outlives that window, and the build fails rather than serving a stale one.

The same rule, applied to the model

Any tool can tell you what it does when it has an answer. These are the places Reach120 has no answer, and what it shows you instead of one.

  • A new account starts empty, not averageYour mastery estimates are built from your own scored responses, so before you have written any there is nothing to build them from. A brand-new account shows an empty list of repeated mistakes rather than the mistakes a typical learner makes. That is the honest state, and it is the state the cold-start audit checks for.
  • An error type with too little evidence is not fitted on its ownEach modelled error type gets its own learn, slip and guess rates only once enough sequences exist to estimate them. The ones below that bar are pooled into a single shared fallback instead of being handed parameters fitted on almost nothing, which would look identical on screen and mean nothing.
  • The forecast declines to fit rather than extrapolateThe filter behind the band forecast needs a minimum number of scored sessions before it will estimate a level and a direction at all. Below that it returns nothing, and the section reports insufficient history instead of drawing a trajectory through one or two points. Typical error is about half a band even when it does fit, so it is a study signal and never a predicted TOEFL score.
  • A section without a forecast is structurally unable to show a bandThis is the part that is not a matter of care. Insufficient history is a state in the API contract, and the contract rejects any response in that state that carries a band, a score, a readiness label or an activity date. A bug that tried to fill the gap with a plausible number would fail validation rather than reach you — the refusal is enforced by the schema, not by remembering to be careful.
  • Where the router has no evidence, it says which kind of answer you gotAdaptive routing that updates from observed outcomes runs in Writing and Listening. Reading and Speaking are routed by deterministic rules over your own weakest skills — still chosen from your evidence, but not adaptive in the same sense. Every recommendation carries the kind of decision that produced it, so the two are distinguishable in the response itself rather than blended into one confident-sounding suggestion.
  • What the routing has learned, it learned from everyoneThe arm estimates behind the router are population-level — one shared set, not a model fitted to you alone. What is yours is the weighting toward your own weakest skills and the draw itself, which is what makes two learners on the same band diverge. Nobody should read the routing as a model trained on their data alone, and we do not describe it that way.
  • A write-up that does not match the computed facts is thrown awayWhen a language model is used to put your figures into sentences, it is handed the finished values and allowed to narrate them and nothing else. Its output is checked back against those values first: a sentence containing a number, a pattern, a route or a claim about your own practice history that the engines did not produce fails the check, and the write-up is discarded rather than corrected or shown. When that happens you get no prose, which is the intended outcome.

The rules above govern the figures this site publishes about itself. The same rule governs the model that runs inside it: where the evidence does not support an answer, the honest output is no answer, and it is enforced in code rather than left to judgment. The methods these behaviours belong to are named, with their measured limits published beside them, on how Reach120 works.

The figures published about Reach120

Each one below is a complete count of the rows that qualify, with company accounts and automated test accounts excluded.

The people who use it, counted by a stated rule

  • 897learners have signed upRecomputed live
    One unit of this figure is
    learner accounts registered on Reach120, excluding admin, company-domain and harness-created accounts.
    How it is computed
    Production database, public_measured_stats() → learners. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    What it leaves out
    89 accounts were excluded before this figure was counted — 57 company-domain accounts, 22 accounts created by a test or factory harness, 10 admin accounts. They are excluded from the metric only; none of them is deleted, and the rule that classifies them is public in scripts/db/2026-08-11-test-identity-marking.sql.
    Measured
    September 19, 2026 (UTC)
  • 466of them joined in the last 90 daysRecomputed live
    One unit of this figure is
    learner accounts registered in the last 90 days, same exclusions.
    How it is computed
    Production database, public_measured_stats() → learners_joined_90d. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    What it leaves out
    89 accounts were excluded before this figure was counted — 57 company-domain accounts, 22 accounts created by a test or factory harness, 10 admin accounts. They are excluded from the metric only; none of them is deleted, and the rule that classifies them is public in scripts/db/2026-08-11-test-identity-marking.sql.
    Measured
    September 19, 2026 (UTC)
  • 5,517Writing and Speaking responses scoredRecomputed live
    One unit of this figure is
    Writing and Speaking responses that have been scored by AI, excluding responses from internal and harness-created accounts.
    How it is computed
    Production database, public_measured_stats() → ai_scored_responses. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    What it leaves out
    89 accounts were excluded before this figure was counted — 57 company-domain accounts, 22 accounts created by a test or factory harness, 10 admin accounts. They are excluded from the metric only; none of them is deleted, and the rule that classifies them is public in scripts/db/2026-08-11-test-identity-marking.sql.
    Measured
    September 19, 2026 (UTC)
  • 123countries it has been found from in Google SearchDated snapshot
    One unit of this figure is
    countries from which at least one person clicked through to the site from Google Search.
    How it is computed
    Google Search Console, property sc-domain:writing30.com (the domain that still carries the organic traffic for this product), countries with >= 1 click. A count over a fixed window, recorded with its measurement date. It stops being published once it outlives that window.
    Measured
    August 4, 2026 (UTC) · covering the 90 days to 2026-08-04

The practice inventory, counted by a stated rule

How much practice exists, rather than how many people use it. Every figure counts only what can actually be served today — retired items kept as anchors for past sessions are excluded, which is why these run lower than the raw table sizes.

  • 3,468practice items across the six question banksRecomputed live
    One unit of this figure is
    servable practice items summed across six banks — Build a Sentence, Reading, Listening, Speaking, Write an Email and Academic Discussion.
    How it is computed
    Production database, public_measured_stats() → practice_items, summed in SQL from the six per-bank counts published beside it. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 1,137Build a Sentence itemsRecomputed live
    One unit of this figure is
    Build a Sentence items that can be served — active bank rows plus the seeded practice file, excluding retired rows kept only as anchors for past sessions.
    How it is computed
    Production database, public_measured_stats() → build_sentence_items (is_active rows plus the bas-% batch that data/build-a-sentence.json serves). Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 1,312Reading questionsRecomputed live
    One unit of this figure is
    Reading questions in the bank, reachable in practice or inside a mock.
    How it is computed
    Production database, public_measured_stats() → reading_items. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 236Reading passagesRecomputed live
    One unit of this figure is
    Reading passages the questions above are attached to.
    How it is computed
    Production database, public_measured_stats() → reading_passages. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 639Listening questionsRecomputed live
    One unit of this figure is
    Listening questions in the bank, reachable in practice or inside a mock.
    How it is computed
    Production database, public_measured_stats() → listening_items. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 230Listening recordingsRecomputed live
    One unit of this figure is
    Listening recordings the questions above are attached to.
    How it is computed
    Production database, public_measured_stats() → listening_recordings. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 251Speaking promptsRecomputed live
    One unit of this figure is
    active Speaking items across Listen & Repeat and Take an Interview.
    How it is computed
    Production database, public_measured_stats() → speaking_items (is_active only). Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 58Write an Email promptsRecomputed live
    One unit of this figure is
    active Write an Email prompts.
    How it is computed
    Production database, public_measured_stats() → email_prompts (is_active only). Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 71Academic Discussion promptsRecomputed live
    One unit of this figure is
    active Academic Discussion prompts.
    How it is computed
    Production database, public_measured_stats() → discussion_prompts (is_active only). Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 1,055vocabulary words in the published decksRecomputed live
    One unit of this figure is
    distinct words carried by the published vocabulary decks.
    How it is computed
    Production database, public_measured_stats() → vocabulary_words (distinct lexemes in decks where is_published). Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)
  • 2,460example sentences from our own passagesRecomputed live
    One unit of this figure is
    example sentences attached to those words, each taken from a Reach120 passage, stimulus or item rather than written for the deck.
    How it is computed
    Production database, public_measured_stats() → vocabulary_contexts, restricted to the words above. Recomputed on request. Every qualifying row is counted. Nothing is sampled, estimated or extrapolated, and the count is recomputed when this page is served.
    Measured
    September 19, 2026 (UTC)

Outcome evidence: real test results learners report

Learners whose booked test date has passed can tell us what they actually scored, and separately choose whether that result may be counted in a published figure. Nobody sees an official score report: every result here is a number a learner typed, and it stays labelled as reported.

0reported results that may be counted in a published figure — consented, not withdrawn
0reported results stored in total, including those kept private to the learner
3learners who have been asked, after a booked test date passed
0of them answered, counting “not now” and “never ask again”

Not enough data yet — nothing is published from this sample

No average, agreement figure or outcome claim is computed until at least 30 consented results have been reported. None have been reported so far. That floor is a convention rather than a guarantee: below a few dozen self-selected reports, an average describes who chose to answer more than it describes how anyone performed, and no arithmetic repairs that.

What this sample is not yet enough for

  • How closely a Reach120 practice band matched the official bandConsented reported results paired with the practice responses scored before the same sitting, broken down by the model that scored them and reported with the sample size behind each cell.
  • Where the practice estimate runs high and where it runs lowThe same pairs, split by band, so the failure modes are published alongside the agreement rather than averaged into it.

Counted September 19, 2026 (UTC). Accounts belonging to the company and to its automated test runs are excluded from every number above.

The reader-facing version of this section — practice volumes beside the improvement figure that stays withheld — is on outcomes. The same discipline applied to the test itself rather than to us is on TOEFL myths, checked against ETS.

How accurate the scoring is, per model

Every score produced since 2026-08-04 records the model that produced it. Paired with the results learners report afterwards, that is how closely a practice band matched an official one — per model, with the sample size behind each figure and the bias the scorer is already known to have.

Not enough data yet — no accuracy figure is published for any engine

A pair needs two things at once: a practice response stamped with the engine that scored it, and a reported official result from that same learner, consented to being counted. There are 798 responses carrying that stamp and 0 consented results, which pair into nothing yet. Until they do, no agreement figure exists to publish and none is estimated.

Responses stamped with the engine that scored them

  • Scoring engine G · Writing: 255 responses from 99 learners.
  • Scoring engine E · Writing: 228 responses from 20 learners.
  • Scoring engine F · Speaking: 101 responses from 49 learners.
  • Scoring engine C · Writing: 88 responses from 19 learners.
  • Scoring engine A · Writing: 62 responses from 9 learners.
  • Scoring engine D · Speaking: 55 responses from 8 learners.
  • Scoring engine B · Speaking: 9 responses from 4 learners.

Engines are shown under stable anonymised labels: which vendor and model each label maps to is routing we do not publish, but the exact engine is recorded on every scored response and shown to the learner who owns it. A further 6,306 scored responses record no engine, because the scorer was only stamped onto new responses from 2026-08-04 and older rows were deliberately not backfilled with a guess. Those cannot be attributed to any engine and are excluded from everything above. Reading and Listening never appear here: both are graded against an answer key, so no engine produces their bands. Computed September 19, 2026 (UTC).

Known bias in the practice estimate you get today (measured on the engine that scores free accounts)

Two attempts were made to correct the behaviour below and both were measured, found not to work, and reverted rather than shipped. So this is how the scorer behaves in the product right now, and publishing it is not an admission — it is the point.

The estimate pulls toward the middle of the scale
Weak responses are marked up and strong responses are marked down at the same time, so the spread of estimates is narrower than the spread of real ability. Across all three re-measured runs of the held-out set, every estimate landed between 1.0 and 4.5, and 78% of them between 2.5 and 4.5.
The lower half of the scale is marked high
On the held-out set, responses written to rubric level 2 were estimated about 0.77 of a point high and responses written to level 3 about 0.71 high; levels 1 and 4 about 0.27 high. If you are early in your preparation, the practice estimate is more likely to flatter you than to discourage you — which is the direction that matters most, because it is the one that could send someone into a test they are not ready for.
The top of the scale is marked low, and 5.0 is effectively unreachable
Responses written to rubric level 5 were estimated about 0.56 of a point low, and the prompt as shipped awarded a 5.0 to none of the 40 held-out responses in any of its three runs — including none of its 24 attempts at level 5. A strong writer should read a 4.0 or 4.5 here as compatible with a higher official result, not as a ceiling.
Typical distance from the intended rubric level
Across the held-out set the estimate sat about 0.52 of a rubric point away from the level the response was written to, in either direction — mean absolute error, 0.49, 0.50 and 0.56 across three runs of the prompt as shipped.
What this was measured against
These responses were written to the published ETS descriptors and graded against them by this project, not scored by ETS. So this measures how consistently the scorer applies our reading of the rubric. It is not a measurement of agreement with an official result, and nothing here should be read as one.

Measured August 5, 2026 (UTC). Source: docs/reach120/32-SCORING-MODEL-BENCHMARK.md §12.2 — 40 held-out responses written to the ETS band descriptors, scored three times by the live production prompt as shipped after F-007; raw per-response runs in docs/reach120/benchmark-runs/scoring-accuracy-f007-*.json.

What is deliberately not computed

A figure that is not published is not measured either, so it cannot quietly start being published.

  • How many people have paidNot published and not counted. It depends entirely on what counts as having paid, and a figure whose definition moves is not evidence of anything.
  • Speaking as its own figureScored Speaking responses are counted inside the scored-response total rather than advertised separately.
  • Anything about improvementNo figure here says practice on Reach120 raised anyone’s score. Establishing that needs the outcome sample above, paired with practice history, and it does not exist yet.

If a figure here looks wrong

Tell us and it gets checked against the query that produced it.

Every claim on this site is registered with its evidence and its measurement date, and a claim that fails re-verification is corrected or withdrawn rather than left standing. Contact us with the figure and where you saw it. The Trust Center covers how scoring works and what Reach120 does not claim.

Reach120 is an independent practice tool. It is not affiliated with, endorsed by, or approved by ETS, and it does not provide official TOEFL® test scores.

TOEFL® is a registered trademark of ETS. This product is not endorsed or approved by ETS.