Skip to main content
20% off your first payment$5.99 for your first week, then $7.49 · Ends November 15, 2026See plans

How the TOEFL is scored by AI in 2026, task by task

Almost every response you give on the current test is graded without a person reading it first. That is not a rumor and it is not a scandal — ETS says so itself, names the systems involved, and describes where a human still sits. Here is the actual division of labor, and what it means for the number you get back from any practice tool.

Written by the Reach120 research teamReviewed
Start free practice

The short answer

Yes, the TOEFL is scored by machines and by AI, and it has been for longer than the current wave of prep marketing suggests. What changed in 2026 is that the scoring is described openly, the report scale is different, and the practice products around the test have started using the same words for very different things.

  1. 1Items with a single right answer are graded against a key by rule-based software. No language model is involved and none needs to be.
  2. 2Open responses — the two longer writing tasks and the Speaking tasks — are graded by an AI scoring engine, because there is no key to match against.
  3. 3On the real test, that engine is not the last word: ETS describes a human in the loop before an official result is issued.
  4. 4A practice tool is a fourth thing entirely. It grades your practice, not your test, and nothing it produces is submitted anywhere.

Which tasks are machine-graded and which are AI-graded

The 2026 specification is explicit about this, and it does not split the way most prep pages assume. "Machine-scored" is not a synonym for "multiple choice", and one of the three writing tasks sits on the machine side.

What ETS states

  • Reading and Listening items are machine-scored selected responses; the Write an Email, Write for an Academic Discussion, and all Speaking tasks are AI-scored constructed responses, per the specification’s item-type notes.

    Build a Sentence is also machine-scored (the Writing section table lists 10 machine-scored items and 0 AI-scored), so "machine-scored" is not exclusive to Reading and Listening. The specification is subject to minor revisions until the official launch.

  • As of January 21, 2026, the TOEFL iBT Writing section consists of three task types: Build a Sentence, Write an Email, and Write for an Academic Discussion.

  • The 2026 specification lists the Speaking section at 11 items and 55 raw points, with task types Listen and Repeat (7 items) and Take an Interview (4 items).

Sourced from the ETS 2026 test specifications, read (link confirmed ). Everything outside this box is Reach120’s own reading of it.

TaskHow it is gradedWhat that means for you
Reading itemsRule-based, against a keyRight or wrong. Nothing to appeal to and nothing a model can misread.
Listening itemsRule-based, against a keySame. Your comprehension is tested; your phrasing never is.
Build a SentenceMachine-scored, per the Writing section tableA word-order task with a correct arrangement, so it is graded like a selected response rather than an essay.
Write an EmailAI-scored constructed responseJudged on the rubric, not against a model answer. Two different good emails can both do well.
Write for an Academic DiscussionAI-scored constructed responseThe same, with the added demand that you engage with what the other speakers said.
Listen and RepeatAI-scored constructed responseYour recorded speech is the input. Intelligibility is what carries it, not accent.
Take an InterviewAI-scored constructed responseExtended spoken answers, graded by engine on the real test with the human step described below.

ETS uses two different systems, and says so

The single most useful thing ETS has published on this subject is that its own practice scores and its own real scores do not come from the same place. Read that carefully before you judge any third party.

What ETS states

  • ETS states that a mock-test score comes "from something called TOEFL AI Lite", which is "built for speed and immediate feedback, helping you identify strengths and weaknesses in real-time", and that "when you take the real test, it is scored by a different system called TOEFL AI Deep".

    Those two names are ETS product names. No prep company, this one included, can use them for its own scoring, and no third party may.

  • On the real-test engine, ETS states it "takes a more comprehensive approach" and "integrates with advanced security protocols with a human in the loop, to ensure your official score is accurate, fair, and recognized by institutions worldwide".

    This is the human step. It is described as part of the real-test pipeline, not as part of the mock-test one.

  • ETS also states that "you might see some small differences between your mock test scores and your real test scores, and that’s totally normal."

    The organization that built both systems, with full access to both, tells you their outputs can disagree. A third-party practice tool has strictly less information than that and should be read accordingly.

Sourced from ETS's TOEFL blog, “Why TOEFL Uses Two AIs”, read . Everything outside this box is Reach120’s own reading of it.

So the honest hierarchy runs: the real-test engine with its human step produces the result universities read; ETS’s own mock engine produces a fast approximation of it and is allowed to differ; and everything outside ETS is a further step removed again.

What the score itself looks like now

What ETS states

  • For TOEFL tests taken on or after January 21, 2026, official score reports use a 1–6 scale in half-band increments.

    Any tool still handing you a 0–120 total as your score is describing the retired report. The 0–120 number now exists only as a comparable estimate ETS provides during its transition through January 2028.

Sourced from the ETS 2026 test specifications, read (link confirmed ). Everything outside this box is Reach120’s own reading of it.

This matters for scoring specifically, because a half-band is a coarse unit. The distance between 4.0 and 4.5 covers a lot of ground, which is both why automated grading can be consistent at this resolution and why an estimate that lands one step out is not a scandal — it is the size of the smallest step that exists.

  • The 1–6 scale in full

    What each band is, how the half-band increments work, and how the 0–120 estimate relates to it during the transition.

  • What each band means

    A page per band, with what it looks like in practice and what it tends to clear.

  • Score calculator

    Convert between the bands, CEFR and the previous scale. No account.

How a practice tool grades the same response — and how ours does

A practice product cannot reach the engine that will grade your test. What it can do is apply the published rubric to your response and be honest about which parts of that are computed and which are a model’s judgment. Reach120’s split is below, stated as narrowly as it is true.

StepWhat happensThe limit it carries
Selected responsesReading and Listening answers are graded against validated answer keys, held server-side so the key never travels with the question.The validation is automated, not a subject-matter review of every question — an automated check can prove a key is consistent and well-formed, and cannot prove it is the right answer.
Open responsesOn an open response, the 0-5 rubric estimate is a judgment an AI model makes; every mastery value, readiness figure and routing decision after it is worked out in code, and the language model only puts those into words. A narration containing a number the engines did not produce is rejected rather than shown.This describes how the pieces are wired. It is not a claim that the underlying estimate is right.
Repeated-error trackingBayesian Knowledge Tracing across 16 modelled error types, fitted per learner, so a recurring problem separates itself from a one-off slip.A learner with no history is shown no patterns rather than an invented one.
Forward-looking bandA gradient-boosted forecast, cross-validated on held-out learners, estimating the band your next practice session is heading for.Typical error is about half a band, so it is a study signal — not a predicted TOEFL score.

Notice what is absent from that table: any figure describing how closely the estimate agrees with a human grader. Reach120’s internal scorer benchmark runs on a corpus too small to publish as an accuracy claim, so it is deliberately not published, and no competitor on this SERP has published a checkable one either.

Four questions that separate real scoring from a scoring claim

  1. 1Which rubric is it applying, and is that rubric the 2026 one? A tool grading an Independent essay is grading a task the test retired.
  2. 2Which part of the number did software compute and which part did a model decide? Both are legitimate; conflating them is not.
  3. 3What is the stated error? A tool that will not name one has not measured one.
  4. 4What does it refuse to tell you? A scoring product with no stated limits has not thought about its limits, or has decided not to mention them.
  • The full method

    Every figure Reach120 publishes about itself, what it counts, where it comes from, and when it was measured.

  • Trust Center

    What is claimed, what is refused, and the independence statement in full.

  • TOEFL and AI: the whole category

    The three different things called "TOEFL AI", and which of them produces a score anyone will read.

  • Try it on your own writing

    Paste a response and see the rubric-shaped read, with no account.

TOEFL AI scoring: common questions

Is the TOEFL scored by AI or by a human?
Both, at different stages and for different task types. Selected responses are graded by rule-based software against a key. Open responses — Write an Email, Write for an Academic Discussion and the Speaking tasks — are AI-scored constructed responses per the 2026 specification. For the real test, ETS describes its engine as integrating "with advanced security protocols with a human in the loop" before an official result is issued.
Which TOEFL tasks are graded without a person reading them?
On the current format, all of them are graded automatically in the first instance. Reading items, Listening items and Build a Sentence are machine-scored against keys. Write an Email, Write for an Academic Discussion, Listen and Repeat and Take an Interview are AI-scored. The human step ETS describes sits in the real-test pipeline around that engine, not in place of it.
Does ETS use the same AI for practice tests and the real test?
No, and it says so plainly. ETS states that mock-test scores come from a system it calls TOEFL AI Lite, "built for speed and immediate feedback", while the real test "is scored by a different system called TOEFL AI Deep". It also states that small differences between the two are normal. Those are ETS product names; no prep company may use them.
Can AI scoring be wrong on the TOEFL?
ETS acknowledges variation between its own mock and real results, which is the closest thing to a published error statement anyone in this space offers. The formal route if you disagree with an official result is ETS’s own score-review process, not a prep tool. A practice estimate from any third party, including this one, has no standing in that process at all.
How accurate is AI scoring on TOEFL practice tools?
Nobody in this category has published a scorer-versus-human agreement figure you could check, and that includes Reach120 — the internal benchmark runs on a corpus too small to carry a published accuracy claim. What is published, on the product page, is the band forecast’s cross-validated error against naive baselines, with its limit stated: typical error is about half a band, which makes it a study signal rather than a predicted score.
Why does Build a Sentence not get AI-scored?
Because it has a correct answer. The Writing section table in the 2026 specification lists 10 machine-scored items and 0 AI-scored ones, which is the specification describing Build a Sentence as a task graded against a key. Routing a task like that through a language model adds a source of error and buys nothing.
Does an AI score mean my accent is being judged?
The Speaking tasks are graded on whether you can be understood, not on sounding like a particular variety of English. Reach120 describes this as intelligibility for that reason, and the free Listen and Repeat drill reads your recording against the target text rather than against an accent model.

Keep going

Reach120 is an independent practice tool. It is not affiliated with, endorsed by, or approved by ETS, and it does not provide official TOEFL® test scores.

TOEFL® is a registered trademark of ETS. This product is not endorsed or approved by ETS.