Skip to main content
20% off your first payment$5.99 for your first week, then $7.49 · Ends November 15, 2026See plans

AI feedback on your TOEFL practice: what to act on, and what to ignore

Every AI prep tool now returns a band and a paragraph of advice on your writing and speaking. The band is the part people read and the least reliable part of the report; the specific, checkable observations underneath it are where the value is. This page explains how to tell them apart — on any tool, including this one.

Written by the Reach120 research teamReviewed
Get feedback on a response

What is actually in an AI feedback report

Strip the presentation off and almost every report in this category contains the same four layers. They are not equally trustworthy, and the order below is roughly the order of how much weight each one deserves.

LayerWhat it isHow much to trust it
Mechanical correctionsGrammar, spelling, word form, punctuation, article and preposition errors, usually with the fix shown inline.High. This is a solved problem, it is checkable in seconds, and a wrong correction is obvious once you look at it.
Task-response observationsWhether you answered the prompt, engaged with the other speakers, covered both parts of a two-part question, or ran out before finishing.High, because these are facts about your text rather than judgments. Verify by rereading the prompt.
Rubric commentaryProse about development, organization, range of language, cohesion — the things the rubric names.Moderate. Usually directionally right, often generic. Useful when it points at a specific sentence, weak when it does not.
The band estimateA single number attached to the whole response.Lowest, and the most read. It is one system’s judgment on a coarse scale, from a tool with no access to the engine that will grade your test.

What good feedback looks like on each 2026 task

The current test has three writing tasks and two speaking tasks, and the useful feedback is different on each. A tool that returns the same generic essay commentary regardless of task is not reading the task.

What ETS states

  • As of January 21, 2026, the TOEFL iBT Writing section consists of three task types: Build a Sentence, Write an Email, and Write for an Academic Discussion.

  • The 2026 specification lists the Speaking section at 11 items and 55 raw points, with task types Listen and Repeat (7 items) and Take an Interview (4 items).

Sourced from the ETS 2026 test specifications, read (link confirmed ). Everything outside this box is Reach120’s own reading of it.

TaskFeedback that helpsFeedback that is noise
Build a SentenceThe correct arrangement and the grammatical reason it is correct — word order, agreement, the function of the connector.A band and a paragraph of prose. This task has a right answer; commentary is a substitute for telling you what it was.
Write an EmailWhether the register matches the recipient, whether every requested point was covered, and whether the purpose is clear in the first line.Praise for "good structure" on a text that never states what it is asking for.
Write for an Academic DiscussionWhether you actually engaged with the named speakers rather than writing past them, and whether your position is stated and supported.Generic essay advice imported from the retired Independent task, which had no other speakers to engage with.
Listen and RepeatA read of your recording against the target text — which words came through and which did not.A judgment about your accent. The question is whether you were understood.
Take an InterviewWhether the answer was complete, on-topic and long enough for the time given, plus specific spots where meaning broke down.A fluency figure with no definition attached to it.

Why the band is the softest part of the report

Three separate reasons stack up, and none of them is a criticism of any particular product.

  • The scale is coarse. Official reports use half-band increments on a 1–6 scale, so the smallest step available is wide, and a response near a boundary can fall either side of it on any given read.
  • The grader is not the grader. No third-party tool can reach the engine that will score your test, so its estimate is a parallel judgment rather than a preview of one.
  • One response is a small sample. A band from a single essay written on a bad afternoon is a data point, not a level.

What ETS states

  • ETS states that "you might see some small differences between your mock test scores and your real test scores, and that’s totally normal."

    ETS built both systems and has access to both. If its own two numbers can disagree, a third party’s estimate carries at least that much spread.

  • For TOEFL tests taken on or after January 21, 2026, official score reports use a 1–6 scale in half-band increments.

    A tool still returning a 0–120 total as your band is reporting on the retired scale.

Sourced from ETS's TOEFL blog, “Why TOEFL Uses Two AIs”, read . Everything outside this box is Reach120’s own reading of it.

What Reach120 does with the part it cannot compute

The interesting engineering question in this category is not how to produce feedback — one model call does that — but what to do about the fact that one step in the chain is a model’s opinion. The answer here is to isolate that step and refuse to let anything downstream invent around it.

  • On an open response, the 0-5 rubric estimate is a judgment an AI model makes; every mastery value, readiness figure and routing decision after it is worked out in code, and the language model only puts those into words. A narration containing a number the engines did not produce is rejected rather than shown.
  • Repeated errors are tracked with Bayesian Knowledge Tracing across 16 modelled error types, fitted per learner, so the report can separate a habit from a slip — and a learner with no history is shown no patterns rather than an invented one.
  • Where there is not enough evidence to compute a figure at all, the report shows nothing rather than an estimate. The internal contract rejects a response that reports insufficient history while still carrying a band.
  • Reading and Listening feedback takes a different path entirely: answers are graded against validated answer keys, and the validation there is automated, not a subject-matter review of every question.
  • The full method

    What is deterministic, what the model decides, and what remains unproven.

  • Sample scored responses

    Real practice responses with the estimated band and the reasoning shown, so you can judge the feedback before writing anything.

How to actually use the feedback you get

  1. 1Fix the mechanical corrections first, by hand, without pasting the text back in for a rewrite. The rewrite teaches you nothing; the correction does.
  2. 2Reread the prompt with the task-response notes beside you and mark anything you did not cover. This is where bands are actually lost.
  3. 3Pick one rubric observation and rewrite a single paragraph against it. One targeted rewrite beats a whole new essay.
  4. 4Ignore the band on any individual response. Look at it across five or ten and only then treat the direction as real.
  5. 5Write the next response before reading old feedback again. Feedback you cannot recall unprompted has not been learned.

Where to get feedback without paying

Reading and Listening practice is free and carries no cap. Writing and Speaking share a free AI-scored response each day. The published material — questions, answers, transcripts, sample responses, PDF packs — needs no account at all.

AI feedback on TOEFL practice: common questions

Is AI feedback on TOEFL writing worth using?
Yes, for the concrete parts. Mechanical corrections and task-response observations are checkable, immediate and genuinely useful. The rubric prose is worth reading as a hint about where to look. The band estimate attached to a single response is the least reliable part of the report and the part most people read first.
Can AI tell me why I lost points on a TOEFL task?
It can tell you what is wrong with the response — an unaddressed part of the prompt, a register mismatch in an email, a position never actually stated in an academic discussion. It cannot tell you what a grading engine it has no access to would have deducted, and a tool phrasing it that way is overstating what it knows.
Is AI feedback the same as an AI score?
No, and the difference matters. Scoring is the pipeline that produces a number. Feedback is what arrives alongside it. A tool can give useful feedback with a shaky number, and a confident number with useless feedback. Judge them separately.
Does AI feedback work for TOEFL Speaking?
For the 2026 speaking tasks it depends on whether the tool can handle audio at all — a general chatbot cannot play or record it. Where it works, the useful output is a read of your recording against the target text and specific spots where meaning broke down. Intelligibility is what the section is about, not accent.
Should I rewrite my essay using the AI’s suggestions?
Apply corrections yourself rather than accepting a rewrite. A rewritten essay is the model’s writing, so it produces a better-looking text and no learning. Fixing the errors by hand, then rewriting one paragraph against one rubric observation, is what changes how you write the next one.
How much AI feedback is free on Reach120?
Reading and Listening practice is free and uncapped. Writing and Speaking share a free AI-scored response each day, and the published questions, answers, transcripts and sample responses need no account at all. Paid passes cover the AI-scored volume beyond that, which is the part with a real per-response cost.
Can AI feedback replace a tutor?
No, and Reach120 will not claim it does. Feedback on one response is not the same as a person who knows your history and what you are avoiding. Tutors use Reach120 to assign practice and read the results, which is the honest shape of the relationship.

Keep going

Reach120 is an independent practice tool. It is not affiliated with, endorsed by, or approved by ETS, and it does not provide official TOEFL® test scores.

TOEFL® is a registered trademark of ETS. This product is not endorsed or approved by ETS.