Correction, September 2026: an earlier version of this page reported agreement figures from a 1,200-essay comparison with human examiners. We cannot substantiate those figures, so we have removed them, along with every place on the site that quoted them.
This page now sets out what is actually known, and what we do and do not measure.
The question behind this page is a fair one: how close is an AI band score to the band a real IELTS examiner would give? The honest answer starts with how the real marking works.
How human examiners mark IELTS Writing
IELTS Writing is marked by trained, certificated examiners against four public criteria, each worth a quarter of the task score: Task Achievement (Task 1) or Task Response (Task 2), Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy.
Each criterion has a published band descriptor, and Task 2 counts for twice as much as Task 1 in the Writing band. Examiners are standardised and monitored, and scripts can be second-marked. The full criteria are published on IELTS.org.
That process is the reference point. No practice tool, AI or human, can reproduce it outside the test. Anything you get before test day is an estimate of that judgement.
What an AI grader actually does
A language model reads your essay alongside the band descriptors and produces a band for each criterion and a list of comments. Done carefully, that gives you three things a human tutor often cannot give on demand:
- Speed. Feedback in seconds, so you can rewrite the same paragraph five times in an evening.
- Criterion-level detail. A separate band for each of the four criteria shows which one is capping your score.
- Pattern spotting. Repeated linkers, missing articles or a vague overview are easy for a model to flag every time they appear.
It also has limits you should know about before you trust the number:
- Bands can drift. A general chatbot asked the same question twice can return different bands for the same essay. A single number from a single run proves little.
- Models can be generous or harsh at the edges. Very strong and very weak essays are where descriptor wording is hardest to apply, for people and models alike.
- Models can invent problems. A comment that does not point to a real sentence in your essay is not feedback. It is noise.
How IELTSbiz checks its own marking
We cannot make an AI grader into an examiner. What we can do is make its behaviour checkable, and hold it to rules that are tested in code on every release:
- Four criteria, marked separately. Every Task 1 and Task 2 answer is marked on the four official criteria, and an automated test fails the release if any task type loses one.
- The overall band is arithmetic, not opinion. The overall band is recomputed in code as the equal-weight average of the four criterion bands, using IELTS rounding. The model's own "overall" number is ignored.
- Every correction quotes you. Each error must quote the exact words from your essay, and the report underlines that text in place. If the quote is not in your essay, nothing is underlined, so invented faults are visible.
- Official length rules. The 150-word and 250-word minimums are stated to the marker and checked against your actual word count.
- Honest labels. Writing and Speaking bands are shown as estimates. Where the Speaking marker never heard your voice, pronunciation is labelled as estimated rather than presented as a measurement.
Reading and Listening are different. They have answer keys, so IELTSbiz grades them deterministically, the same way every time, and converts your raw score with a band table. There is no AI judgement in those scores at all.
What we have not measured
We have not run a study comparing our Writing bands with bands from certified IELTS examiners on the same essays. Until we do, and publish the method alongside the result, we will not claim a match rate.
If a site, including a competitor, quotes an agreement percentage, ask how many essays, who marked them, and whether the markers saw the AI score first. A number without that method is a marketing figure, not evidence.
How to use an AI band score sensibly
- Watch the criterion, not the headline. If Lexical Resource is consistently half a band below the others, that is where your study time should go.
- Check every comment against your text. Good feedback points at a sentence you wrote. Discard any that does not.
- Compare like with like. Track the trend across essays marked the same way. Do not compare a band from one tool with a band from another.
- Get a human read before a high-stakes test. If a visa or offer depends on the result, have a qualified teacher mark one or two full practice essays.
The bottom line
AI marking is a fast, repeatable way to find out which criterion is holding you back and to practise fixing it. It is not a substitute for the examiner, and any tool that claims to be should be able to show you exactly how it knows.
You can compare our criteria with the official IELTS marking criteria and the British Council's guidance on how IELTS is assessed.