What to prepare first
We built a regression set from 17 official IELTS samples with examiner scores and comments, spanning Band 4 to 8.5. These samples provide both a credible score and examiner evidence that can be checked against the model output.
If you already have a real essay, a target band, and exam pressure, find the blocker holding the score back and connect it to a verifiable revision. Do not stop at another AI correction list.
This is a score-gap diagnostic, not an official IELTS score prediction.
On 17 official IELTS samples, the first rubric with DeepSeek flash scored Band 7–7.5 essays at 4.5–6 and Band 8–8.5 essays at 5–7. Mean absolute error was 1.22–1.39, and no prediction exceeded 7.
Reliable feedback cannot stop at a realistic-looking score. It cites the essay, maps evidence to criteria, selects the main blocker, and gives a verifiable next-version revision step.
First of all, technology saves a lot of time in our daily life. For example, when we go to a new place, we do not need to ask people or remember the way, we just use the map App on the phone. Besides, we can pay by phone, buy things online and talk with friend…
Marked line: where the diagnosis located the main blocker in this learner-written essay.Coherence & cohesion 6
Sample diagnostic, not an official IELTS score.
We built a regression set from 17 official IELTS samples with examiner scores and comments, spanning Band 4 to 8.5. These samples provide both a credible score and examiner evidence that can be checked against the model output.
The ceiling came from stacked biases: conservative scoring, counting absolute errors rather than frequency, penalising one defect across several criteria, missing high-band anchors, and model differences. Five errors in 350 words are not equivalent to five in 180, and one expression problem should not be punished four times.
The v3 rubric added anchors through Band 9, allowed occasional high-band errors, removed blanket downward scoring, mapped each criterion to official descriptors, and used official half-band rounding. MAE narrowed to 0.83–1.00, yet many Band 7.5–8.5 Task 2 samples still scored 5.5–6.5. A better rubric could not fully compensate for model judgement.
After cross-model comparison, the system moved to Qwen. From 16–25 July 2026, ten consecutive five-case scheduled regressions passed with MAE at 0.3–0.4; 35 fixed-sample cases ran in the most recent seven days. This shows improved stability on those calibration points and much less mechanical compression around 6–6.5.
A scheduled regression on five fixed samples is not a new independent validation. The full v7 17×3 engineering gate was stopped after 60 minutes without a final report, and a fully untouched blind set has not yet been completed. We therefore do not claim proven accuracy, superiority to examiners, an accuracy percentage, or guaranteed improvement.
For a learner stuck between 6.5 and 7, an isolated number is not enough. Useful AI feedback identifies the criterion holding the score back, quotes the relevant essay evidence, and gives a revision that can verify change. With ChatGPT or any other tool, provide the full prompt, essay, target band, and concern—then treat the output as clues rather than a verdict.
If you already have a real essay, a target band, and exam pressure, find the blocker holding the score back and connect it to a verifiable revision. Do not stop at another AI correction list.
On 17 official IELTS samples, the first rubric with DeepSeek flash scored Band 7–7.5 essays at 4.5–6 and Band 8–8.5 essays at 5–7. Mean absolute error was 1.22–1.39, and no prediction exceeded 7.
Reliable feedback cannot stop at a realistic-looking score. It cites the essay, maps evidence to criteria, selects the main blocker, and gives a verifiable next-version revision step. The ceiling came from stacked biases: conservative scoring, counting absolute errors rather than frequency, penalising one defect across several criteria, missing high-band anchors, and model differences. Five errors in 350 words are not equivalent to five in 180, and one expression problem should not be punished four times.
Banfen Writing focuses on targeted score-gap repair. Use a real essay to pinpoint your conservative band range, core blockers, text evidence, and next revision steps.
These pages explain the Banfen Writing diagnostic and blocker-repair method, Task 1 overview, Task 2 development, and common score blockers.
Submit an essay you just wrote to see a conservative range, the priority problem, and evidence from your draft. Once signed in, use a free Repair Cycle to carry the repair into your own full second draft.
More than a simple comment list. Identify the priority blocker in your draft and guide you to revise with proven results.