Skip to main content

TOEFL and AI in 2026: what is scored by a machine, and what a practice tool can actually tell you

Three different things get called "TOEFL AI": the engines ETS uses to score the real test, the AI inside practice products, and general chatbots people prepare with. They are not interchangeable, and only one of them produces a score any university will read.

Written by the Reach120 research teamReviewed
Start free practice

The short answer: three different things share the name

If you searched for "TOEFL AI" you were probably after one of three things, and the answer is different for each.

What you might meanWhat it actually isDoes it produce an official score?
ETS's own scoring AIThe engines ETS uses to score responses on its mock tests and on the real testThe real-test engine does. It is the official score.
An AI practice toolA third-party product that scores or comments on practice responses, Reach120 includedNo. It produces a practice estimate, never a score anyone can submit.
A general chatbotA general-purpose assistant used to draft, critique, or drill answersNo, and it cannot deliver the audio tasks the 2026 test contains.

What ETS means by AI on the TOEFL test

ETS is explicit that the current test is scored by machines and by AI, and that its practice products and its real test do not use the same system. It has named both.

What ETS states

  • ETS states that a mock-test score comes "from something called TOEFL AI Lite", which is "built for speed and immediate feedback, helping you identify strengths and weaknesses in real-time", and that "when you take the real test, it is scored by a different system called TOEFL AI Deep".

    Those two names are ETS product names. No prep company, this one included, can use them for its own scoring.

  • On the real-test engine, ETS states it "takes a more comprehensive approach" and "integrates with advanced security protocols with a human in the loop, to ensure your official score is accurate, fair, and recognized by institutions worldwide".

  • ETS also states that "you might see some small differences between your mock test scores and your real test scores, and that’s totally normal."

    Worth reading twice: ETS says its own practice scores and its own official scores can differ. No practice product of any kind is exempt from that.

Sourced from ETS's TOEFL blog, “Why TOEFL Uses Two AIs”, read August 6, 2026. Everything outside this box is Reach120’s own reading of it.

The 2026 specification is more precise about which task types are handled which way, and this is the part most prep pages get wrong.

What ETS states

  • Reading and Listening items are machine-scored selected responses; the Write an Email, Write for an Academic Discussion, and all Speaking tasks are AI-scored constructed responses, per the specification’s item-type notes.

    Build a Sentence is also machine-scored (the Writing section table lists 10 machine-scored items and 0 AI-scored), so "machine-scored" is not exclusive to Reading and Listening. The specification is subject to minor revisions until the official launch.

  • As of January 21, 2026, the TOEFL iBT Writing section consists of three task types: Build a Sentence, Write an Email, and Write for an Academic Discussion.

  • For TOEFL tests taken on or after January 21, 2026, official score reports use a 1–6 scale in half-band increments.

    Any tool still handing you a 0–120 total as your score is describing the retired report. The 0–120 number now exists only as a comparable estimate ETS provides during its transition through January 2028.

Sourced from the ETS 2026 test specifications, read July 27, 2026. Everything outside this box is Reach120’s own reading of it.

What an AI practice tool actually does

Strip the marketing off the category and almost every product does the same four things. Knowing which one you are being sold makes the comparison easy.

  1. 1Delivers the task. It shows you a prompt, plays audio where the task has audio, and holds you to the clock. This is content work, not AI, and it is where most tools are quietly out of date — a product still serving Integrated and Independent essays is drilling a test that no longer exists.
  2. 2Marks what has a right answer. Reading and Listening selected responses, and Build a Sentence, are graded against a key. Nothing about that step needs a language model, and a tool that routes it through one is adding a source of error for no gain.
  3. 3Comments on what has no single right answer. Email, Academic Discussion and the Speaking tasks are open responses, so the tool produces a rubric-shaped estimate and some feedback. This is the AI part, and it is the part with real error bars.
  4. 4Decides what you practise next. Most products use a fixed syllabus or let you pick. A few, this one included, route from your own history.

What no AI practice tool can do — including this one

  • It cannot give you an official score. Official TOEFL scores are issued by ETS, from a test taken under test conditions. Everything else is an estimate for your own tracking.
  • It cannot be calibrated to ETS. No third party has access to ETS’s scoring engines or its rubric weights, so "ETS-calibrated" and "official-level scoring" describe something that is not available to be done. Reach120 does not make that claim and you should distrust anyone who does.
  • It cannot promise band accuracy it has not measured. A tool that publishes an accuracy figure should tell you the corpus, the method and the residual error. One that publishes a round number with no method has published a marketing figure.
  • It cannot tell you your percentile. ETS has not published percentile data for the 1–6 scale. Any percentile you are shown is somebody’s estimate from the concordance, and it should say so.
  • It cannot replace a teacher. Feedback on a response is not the same thing as a person who knows what you did last month and why you keep doing it.
  • A general chatbot has two extra gaps: it cannot play or score the audio tasks that make up the whole Speaking section and part of Listening, and it does not know the 2026 task set unless you paste it in.

How Reach120 works out a practice band

Four engines, and a rule about which one is allowed to speak. The rule is the important part.

EngineWhat it doesThe limit it carries
Glass BoxEvery score, mastery value and recommendation is worked out in code; the language model only puts the result into words. A narration containing a number the engines did not produce is rejected rather than shown.This is an architectural fact about how the pieces are wired, not a claim that the underlying estimate is right.
Mistake MemoryBayesian Knowledge Tracing across 16 modelled error types, fitted per learner, so repeated problems separate themselves from one-off slips.A learner with no history is shown no patterns rather than an invented one.
Band ForecastA gradient-boosted forecast, cross-validated on held-out learners, estimating the band your next practice session is heading for.Typical error is about half a band, so it is a study signal — not a predicted TOEFL score.
Practice RouterThompson-sampling adaptive routing that updates from what actually worked: two students with the same score get different next tasks.It runs in Writing and Listening only — Reading and Speaking use deterministic per-learner rules — and what it learns is population-level, then routed per learner from your own evidence.

Selected responses take a different path entirely. Reading and Listening answers are graded against validated answer keys, held server-side so the key never travels with the question. The validation there is automated, not a subject-matter review of every question — an automated check can prove a key is consistent and well-formed, and cannot prove it is the right answer.

Every free surface, and what each one is for

Reading and Listening practice is free and carries no cap. Writing and Speaking share a free AI-scored response each day. The published material below — questions, answers, transcripts, sample responses, PDF packs — needs no account at all.

  • TOEFL Writing practice

    The hub for all three 2026 writing tasks, each with a drill and a full guide: Build a Sentence, Write an Email, Write for an Academic Discussion.

  • Listen and Repeat practice

    Play a sentence, record your own, get a read against the target text. Sample sets and transcripts are published on the page.

  • Take an Interview practice

    The second 2026 Speaking task, with a published prompt bank and what each part is scored on.

  • Reading practice

    Free and uncapped, including Complete the Words — the 2026 task type almost nobody serves.

  • Listening practice

    Free and uncapped, with the four 2026 listening task types and transcripts published beside the audio.

  • Free practice tests

    Full four-section mocks, scored instantly rather than returned to you in a day or two.

  • Free writing checker

    Paste a response, get a rubric-shaped read with no account. The free path has spend limits, which is why it can stay free.

  • Free tools

    Score calculator, per-section calculators, the TOEFL-to-IELTS converter, and the current-material audit.

  • Downloads

    Practice packs, prompt banks and a complete 2026 practice test as PDFs, ungated.

  • Sample scored responses

    Real practice responses with the band Reach120 estimated and the reasoning shown.

TOEFL and AI: common questions

Does the TOEFL test use AI to score answers?
Yes. Per the 2026 specification, Reading and Listening items and Build a Sentence are machine-scored selected responses, while Write an Email, Write for an Academic Discussion and all Speaking tasks are AI-scored constructed responses. ETS states that its mock tests are scored by a system it calls TOEFL AI Lite and the real test by a different one it calls TOEFL AI Deep, the latter with a human in the loop.
Is "TOEFL AI" a product?
Two of them are, and both belong to ETS: TOEFL AI Lite and TOEFL AI Deep are the names ETS gave its own scoring systems. When a prep company uses the phrase, it is describing a feature, not a product name — Reach120 does not brand anything "TOEFL AI" and no third party may.
Can an AI tool tell me what I would score on the real test?
No. It can estimate a band from your practice response, and a good tool will tell you the size of its typical error. Official scores come only from ETS, from a test taken under test conditions. ETS itself notes that its practice scores and its official scores can differ.
Is AI scoring accurate for TOEFL Writing?
Accurate enough to show you what to fix, not accurate enough to be a promise — and no prep company, this one included, has published a scorer-versus-human agreement figure you could check. Reach120’s own scorer benchmark runs on an internal corpus too small to publish as an accuracy claim, so it is deliberately not published. What is published, on the product page, is the band forecast’s cross-validated error against naive baselines, with its limit stated: typical error is about half a band on the 0–5 rubric, which makes it a study signal rather than a predicted score.
Can I just use ChatGPT to prepare for the TOEFL?
For drafting, brainstorming and grammar questions, a general chatbot is genuinely useful. It cannot deliver the 2026 test’s audio tasks — Listen and Repeat, Take an Interview, and the Listening section — it does not know the current task set unless you paste it in, and its band estimates come with no stated method or error.
What does "ETS-calibrated" mean on a prep website?
Nothing verifiable. No third party has access to ETS’s scoring engines or rubric weights, so calibration against them cannot be performed or checked. Treat the phrase as a marketing claim and ask instead what corpus the tool measured itself against.
Does AI feedback replace a teacher?
No, and Reach120 will not claim it does. Automated feedback is fast, patient and available at 2am; a teacher knows your history, your first language and what you are avoiding. Tutors use Reach120 to assign practice and read the results, which is the honest shape of the relationship.
Is any of this free?
Reading and Listening practice is free and uncapped. Writing and Speaking share a free AI-scored response each day, and the published questions, answers, transcripts, sample responses and PDF packs need no account at all. Paid passes exist for the AI-scored volume beyond that, which is the part that costs real money to run.
How does Reach120 stop the AI from making numbers up?
By not letting it produce them. Scores, mastery values and recommendations are computed in code; the language model only narrates the result, and a narration containing a number, pattern or route the engines did not produce is rejected rather than shown to the learner.

Keep going