Short answer: ChatGPT will give you an IELTS band in about ten seconds, and you should not plan your test date around it. It can read your essay, apply criteria and produce feedback that is genuinely useful.
What it cannot reliably do is produce the same band twice, or a band that survives a change in how you asked the question. The feedback is worth having. The number needs handling with care.
This matters because the band is the part everyone reads. A candidate who is told 7.0 books the test; a candidate told 6.0 delays it and pays for another month of preparation.
If that digit moves depending on whether you told the model to be strict, it is not a measurement — it is a mood.
Why the same essay gets different bands
A language model is not applying a rubric the way an examiner does. It is predicting what a plausible assessment of your text looks like, and "plausible" is shaped by everything else in the conversation: how you framed the request, whether it praised you three messages ago, and what persona you assigned it.
We saw this directly while testing prompts for our AI prompt library.
The same synthetic Band 6.5-ish essay was marked by the same model under two different prompts — one that forces quoted evidence before any judgement, and one that casts the model as a tired, strict examiner who has marked forty scripts that day.
On that run both returned Band 7, which is reassuring but not guaranteed; the strict-persona prompt is designed to suppress flattery, and when it does move a mark, it moves it downward.
The gap between those two numbers on your own essay is a useful diagnostic in itself: it tells you how much the framing, rather than your writing, is driving the score.
| What you change | What tends to happen to the band |
|---|---|
| Praise earlier in the same chat | Drifts up — the model is consistent with its own earlier tone |
| "Be a strict examiner" persona | Drifts down, sometimes past what the essay deserves |
| Descriptors pasted vs recalled from memory | Pasted is materially more stable; recalled wording drifts |
| Evidence demanded before the band | More stable, and the reasoning becomes checkable |
| Asking again in the same chat | Tends to defend the first answer rather than re-examine |
What we measured about AI grading
There is a difference between a general chatbot and a grading system built and calibrated for the task.
We published a double-blind study comparing our own AI grading against three certified and former IELTS examiners across 1,200 real Academic Task 2 essays: bands matched the human consensus within ±0.5 in 94.2% of essays and exactly in 78.3%, with agreement strongest between Band 6.0 and 7.5 — the range most university applicants sit in.
The full method and results are published so you can judge them.
That figure is for a purpose-built grader with a fixed prompt, a fixed model and no conversation history. A general assistant, handed a rubric you pasted and whatever context your chat already contains, is looser than that.
Treating a ChatGPT band as equivalent to a measured one is the single most common mistake we see candidates make with these tools.
The part that is genuinely reliable
Strip away the number and what remains is often excellent. A model reading your essay against the descriptors can tell you which sentence is doing the work of a topic sentence and failing, which pronoun points at nothing, and which paragraph contains two ideas wearing one hat.
Those observations are checkable — you can look at the sentence it quoted and see whether it is right.
The most valuable single output we have found is the answer to why does this essay fail the next band up. It forces a specific claim about a specific gap, and it converts directly into a revision task.
Compare that with "develop your ideas more fully", which is true of nearly every essay and actionable for none of them.
Three things that make any AI mark more trustworthy
- Paste the actual descriptors. Models paraphrase the band descriptors from memory and the paraphrase drifts. The official public descriptors are a free download from ielts.org, and pasting them is the difference between descriptor-anchored marking and an impression.
- Force evidence before judgement. A model made to quote your text before assigning a band cannot work backwards from a number it already chose. This is the single highest-value structural change you can make to a marking prompt.
- Start a new chat every time. Conversation history is the biggest cause of inflated bands. A model that praised your last essay will be reluctant to be hard on this one.
All three are built into the marking prompts in our Writing Task 2 prompt collection, along with the real output each one produced when we ran it and the specific way each one goes wrong.
A workable routine
Use the AI for volume and the measured tool for the number. Write the essay under timed conditions, mark it with an evidence-first prompt, then immediately ask the model to argue the essay is half a band lower — the reasons it produces are your revision list.
When you want a band you can actually plan around, run the essay through a grader whose accuracy has been measured, such as our free Writing Checker, and treat that as the reference point.
The candidates who get the most out of these tools are the ones who stopped asking "what band is this?" and started asking "what specifically is costing me the next half band?" The first question has an unreliable answer. The second has a useful one.