Force the model to quote your essay before it scores it
Marks your Task 2 essay in four locked steps — evidence first, descriptors second, band last — so it cannot work backwards from a flattering number.
Here are the complete official IELTS Writing Task 2 band descriptors: [RUBRIC] Here is my essay: [MY ESSAY] Here is the question: [QUESTION] Mark it in this exact order — do not skip ahead: STEP 1 — Evidence collection only. For each of the four criteria, quote 3-5 passages from my essay that are relevant to that criterion. No judgements yet. STEP 2 — Descriptor matching. For each criterion, quote the descriptor wording at Band 6, 7 and 8, and say which of those three my evidence best matches, and why it fails the next band up. STEP 3 — Bands. One table: | Criterion | Band | The descriptor phrase that decides it | STEP 4 — Overall band with the official rounding applied. Show the arithmetic. STEP 5 — The single highest-value fix, and the band it would realistically produce. If at any point you cannot support a judgement with a quote, say so.
Best on Claude · ChatGPT · Gemini.
What to fill in
| Token | What to paste |
|---|---|
| [RUBRIC] | The official public Writing Task 2 band descriptors, pasted verbatim. Download them free from ielts.org — never let the model recall them from memory. |
| [MY ESSAY] | Your own essay, unedited. Leave the typos in: they are evidence. |
| [QUESTION] | The exact Task 2 question you answered, word for word. |
What you should get back
A four-stage analysis: quoted evidence per criterion, descriptor matching across Bands 6–8 with a reason your essay fails the next band up, a decision table, the rounding arithmetic, and one prioritised fix.
What it actually returned
Tested on Gemini 2.5 Flash, API, August 2026, 2026-08-06. Trimmed — this is an excerpt of a longer response, not the whole thing.
STEP 1 — Evidence collection only. Task Response: 1. "I completely agree with this opinion because of two reasons which I will discuss below." (Clear position) 2. "Firstly, public transport is better for the environment. Cars produce a lot of pollution…" (Extends idea about environmental benefits) 3. "Moreover, cities like Copenhagen have invested heavily in cycling and public transport…" (Supports environmental point with an example) […] STEP 2 — Descriptor matching. Task Response: — Band 6: addresses the task but development is uneven — Band 7: addresses all parts with a clear position and extended ideas — Band 8: sufficiently addresses all parts with well-developed, well-supported ideas Best Match: Band 7. The essay clearly addresses all parts of the question, maintaining a clear position of complete agreement. Ideas are extended (e.g. explaining why public transport is better for the environment, detailing "induced demand"). Why it fails Band 8: While ideas are extended and supported, the development is not consistently well-developed and well-supported to the Band 8 standard. The Copenhagen example is brief…
Our verdict on this run
The forced order held, which is the whole point: Step 1 contains quotes and neutral labels with no judgements, where a one-line "score my essay" prompt would have opened with a band. The "why it fails Band 8" paragraph is the genuinely useful output — naming the Copenhagen example as too brief is specific and actionable in a way that "develop your ideas further" is not. Two weaknesses: the descriptor wording it quoted was only as good as what we pasted in, and it produced the overall band without showing the rounding arithmetic Step 4 explicitly asked for.
How this prompt goes wrong
The model skips ahead. If Step 1 already contains words like "weak" or "effective", it has started judging before it finished collecting, and the band that follows is the one it decided in the first ten seconds. Scroll back: if the evidence section is not purely quotes, discard the run and start a fresh chat rather than asking it to try again — a model that has stated a band defends it. The second failure is subtler: it quotes descriptor wording that does not exist. If a "descriptor phrase" sounds like a paraphrase, check it against the PDF, because a band justified by invented wording is just a guess in formal clothing.
Tips
- The forced order is the whole mechanism. Do not reorder the steps to save time — Step 2 is where the useful paragraph lives.
- Step 2's "why it fails the next band up" is the most valuable output in this prompt. Copy that sentence into your notes; it is your revision brief.
- Run it in a fresh chat each time. Prior praise in the conversation history pulls the band up.
- Compare its band against a measured one from our Writing Checker. If they disagree by more than 0.5, trust the feedback and distrust both numbers.
Related prompts
- Writing Task 2Give the model a known band first, then mark your essayMarks a reference essay whose real band you already know, forces a recalibration, then marks yours against that same standard.
- Writing Task 2Make the model argue your essay is half a band worseAsks the model to build the strongest case against the band it just gave you, then arbitrate honestly between the two.
- Writing Task 2Analyse only coherence and cohesion, nothing elseMaps your controlling ideas, referencing chains and linker repeats — the criterion most learners cannot self-assess — with grammar and vocabulary explicitly off-limits.
- Mock examinerA 250-word verdict with no room for flatteryA tired, strict examiner persona capped at 250 words: the band first, three quoted justifications, one upgrade condition, nothing else.
Common questions
Why does the model still praise me even with this prompt?
Because praise is cheap and the prompt does not forbid it — it only forces evidence first. If flattery persists, add "Never open with praise" as a system or custom instruction, and run the essay in a brand-new chat. Existing conversation history is the biggest single cause of inflated bands.
Do I really have to paste the rubric?
Yes. Models paraphrase the descriptors from memory and the paraphrase drifts, which is exactly the wording Step 2 depends on. Pasting the official PDF text is the difference between descriptor-anchored marking and a vibe.
Can I use this on the free tier?
Yes, on all three recommended tools, though the rubric plus your essay is a long input — a free-tier context limit can truncate it silently. If Step 2 starts inventing descriptor wording, that is usually what happened.
Is the band it gives me reliable?
Treat it as a range, not a number. In our own study of 1,200 essays, purpose-built AI grading matched the human examiner consensus within ±0.5 bands 94.2% of the time — a general chatbot with a pasted rubric is looser than that. The evidence and the "why it fails the next band" reasoning are worth more than the digit.