Give the model a known band first, then mark your essay
Marks a reference essay whose real band you already know, forces a recalibration, then marks yours against that same standard.
You will mark two IELTS Writing Task 2 essays. Here are the official band descriptors: [RUBRIC] ESSAY A (reference — a certified examiner scored this Band [X]): [ANCHOR] Step 1: Mark Essay A against the four criteria. Show your reasoning. If your bands differ from [X], explain what you weighted differently, then recalibrate. Step 2: Now mark ESSAY B using the same standard you just calibrated: [MY ESSAY] For Essay B, output only this table: | Criterion | Band | Two quotes from the essay that justify this band | What would move it up 0.5 | Then: overall band, your confidence (high/medium/low), and the one change with the largest mark impact.
Best on ChatGPT · Claude · Gemini.
What to fill in
| Token | What to paste |
|---|---|
| [RUBRIC] | The official public Writing Task 2 band descriptors, verbatim from ielts.org. |
| [ANCHOR] | One essay whose band you actually know — a teacher's mark, an official Cambridge sample, or a Writing Checker result you trust. |
| [X] | The anchor essay's known band, e.g. 7.0. |
| [MY ESSAY] | Your own essay, unedited. |
What you should get back
A calibration pass on the reference essay (including an admission of what it weighted differently), then a four-row evidence table for your essay, an overall band with a stated confidence level, and one prioritised fix.
What it actually returned
Tested on Gemini 2.5 Flash, API, August 2026, 2026-08-06. Trimmed — this is an excerpt of a longer response, not the whole thing.
Here is the marking for Essay B, following the calibrated standard from Essay A. ESSAY B MARKING | Criterion | Band | Two quotes that justify this band | What would move it up 0.5 | | Task Response | 7 | "I completely agree with this opinion because of two reasons which I will discuss below." (Clear position) / "cities like Copenhagen have invested heavily in cycling and public transport, and as a result their air quality is much better than other cities." (Extends and supports ideas with an example) | Adding a layer of nuance by briefly acknowledging the necessity or specific uses of roads (e.g. for freight or remote areas), even while arguing for public transport's superiority in urban contexts… | | Coherence and Cohesion | 7 | "Firstly, public transport is better for the environment." / "Secondly, building new roads does not solve traffic problems. Moreover, when a new road is built…" | […] |
Our verdict on this run
It skipped Step 1 entirely. We asked it to mark the reference essay first, show its reasoning and recalibrate if it disagreed — and it went straight to the table for Essay B, asserting it had "followed the calibrated standard" without ever showing one. That is the failure this prompt is most prone to and it is easy to miss, because the output that follows looks thorough: the table is well-formed, the quotes are real, and the "move it up 0.5" column names a genuinely useful gap (acknowledging where roads still matter). If you do not see a visible Step 1 with an explicit agreement or disagreement about the anchor band, the calibration did not happen and the marks that follow are ordinary marks.
How this prompt goes wrong
Step 1 silently disappears. In the captured run below the model went straight to marking Essay B, announcing it had "followed the calibrated standard from Essay A" without ever showing a calibration pass — and because the table that follows looks thorough, it is easy to accept. If you cannot see an explicit mark for the anchor with an agreement or disagreement about its band, no calibration happened and what you have is an ordinary mark with a reassuring preamble. The subtler version is compliance: Step 1 appears, returns exactly [X] on all four criteria, and disagrees with nothing. Genuine calibration almost always disagrees somewhere. Test it by running the anchor once with its band withheld; if the blind band is more than 0.5 from the one it "confirmed" when told, the calibration is theatre.
Tips
- Get one real anchor and reuse it forever. Its value is that its band is externally verified, not that it is recent.
- Run it three times in separate chats. If the three bands for your essay spread by more than 0.5, treat all three as unreliable and use the feedback instead.
- Pick an anchor near your target band, not near your current one — the calibration matters most at the boundary you are trying to cross.
Related prompts
- Writing Task 2Force the model to quote your essay before it scores itMarks your Task 2 essay in four locked steps — evidence first, descriptors second, band last — so it cannot work backwards from a flattering number.
- Writing Task 2Make the model argue your essay is half a band worseAsks the model to build the strongest case against the band it just gave you, then arbitrate honestly between the two.
- Writing Task 2Test your introduction and conclusion against seven checksRuns a yes/no scorecard over the two most formulaic paragraphs in Task 2, quotes the evidence, then rewrites both in your own voice.
- Mock examinerFind out what an examiner notices in your first 50 wordsTests your essay the way it is actually read — fast — separating conspicuous errors from invisible ones and naming the feature that tips the band.
Common questions
Where do I get an anchor essay with a known band?
Three reliable sources: an essay a teacher has marked, an official Cambridge sample answer published with its band, or our Task 2 sample library, which publishes Band 9 and Band 6.5 versions of the same question.
Does the anchor have to be on the same topic as my essay?
No, and it is better if it is not. You are calibrating the marking standard, not comparing content. A same-topic anchor tempts the model to compare your ideas against it rather than apply the descriptors.
Why does it always agree with the anchor band?
Because you told it the answer. That is the known weakness of this prompt — see the failure mode above for the blind test that detects it. A calibration pass that never disagrees is not calibrating.