Make the model argue your essay is half a band worse
Asks the model to build the strongest case against the band it just gave you, then arbitrate honestly between the two.
Here is my essay and the band you gave it: [MY ESSAY] — you said Band [X]. Now argue the opposite case. Build the strongest possible argument that this essay is actually Band [X minus 1]. Use the official descriptors: [RUBRIC] Quote my weakest three sentences and explain, in descriptor language, how an examiner having a strict day would justify the lower band. Then tell me honestly: was your first band or this one closer to the truth, and what specifically in the essay makes it ambiguous?
Best on Claude · ChatGPT · Gemini.
What to fill in
| Token | What to paste |
|---|---|
| [MY ESSAY] | The same essay it just marked, in the same chat. |
| [X] | The band it gave you a moment ago. |
| [X minus 1] | One half-band below that — write the actual number (a 7.0 becomes 6.5), because models handle a figure more reliably than the arithmetic. |
| [RUBRIC] | The official public Writing Task 2 band descriptors. |
What you should get back
A prosecution case against your essay with three quoted weak sentences explained in descriptor language, then an honest arbitration naming exactly what makes the essay borderline.
What it actually returned
Tested on Gemini 2.5 Flash, API, August 2026, 2026-08-06. Trimmed — this is an excerpt of a longer response, not the whole thing.
This is a great exercise in applying the descriptors rigorously! Let's argue for a Band 6.0. Argument for Band 6.0 1. Task Response (Band 6: addresses the task but development is uneven; Band 7: addresses all parts with a clear position and extended ideas) — While a clear position is stated and maintained, the development of ideas can be seen as uneven. The first body paragraph makes a general claim about public transport and the environment, then offers a good example (Copenhagen). However, the initial explanation of how public transport specifically reduces pollution beyond "less CO2" is quite brief… — In the second body paragraph, the concept of "induced demand" is introduced effectively, but the subsequent point about roads being "expensive and they take up a lot of space" feels somewhat tacked on with a simple "Also," rather than being deeply integrated…
Our verdict on this run
It opened with praise — "This is a great exercise!" — before doing anything, which is the reflex this whole prompt exists to defeat and a reminder that suppressing it needs a custom instruction, not politeness. Once past that, the prosecution case was real rather than a strawman: it went after the structural weakness in body 2, where a second idea is attached with "Also" instead of developed, and that is exactly the fault a strict examiner would seize on. Note it argued the case at criterion level and never quite delivered the three weakest sentences as quotes that the prompt asked for.
How this prompt goes wrong
The model builds a strawman it can knock down, then reaffirms its original band — a fake concession. The tell is a prosecution case that quotes trivia (one comma, one article) rather than a structural weakness, followed by "however, on balance my original assessment stands". When that happens the essay may genuinely be solidly at that band, but more often the model is protecting its earlier answer. Re-run the essay in a clean chat asking only for the lower-band case with no prior number mentioned; if that case is suddenly persuasive, the arbitration you got was social, not analytical.
Tips
- Run this immediately after any AI band, in the same chat. It is the fastest way to strip out flattery.
- The reasons in the prosecution case are your to-do list. The band itself barely matters.
- Whatever it names as "what makes it ambiguous" is what a real examiner would be deciding on. Fix that first.
Related prompts
- Writing Task 2Force the model to quote your essay before it scores itMarks your Task 2 essay in four locked steps — evidence first, descriptors second, band last — so it cannot work backwards from a flattering number.
- Writing Task 2Give the model a known band first, then mark your essayMarks a reference essay whose real band you already know, forces a recalibration, then marks yours against that same standard.
- Writing Task 2Analyse only coherence and cohesion, nothing elseMaps your controlling ideas, referencing chains and linker repeats — the criterion most learners cannot self-assess — with grammar and vocabulary explicitly off-limits.
- Mock examinerA 250-word verdict with no room for flatteryA tired, strict examiner persona capped at 250 words: the band first, three quoted justifications, one upgrade condition, nothing else.
Common questions
Is my essay really the lower band?
Usually the truth is between the two, which is why the arbitration step exists. The useful output is not the band — it is the list of what makes the essay borderline, because that is precisely what a real examiner would be weighing.
Should I use this on a band I am happy with?
Especially then. An essay you were told is 7.5 that comes back defensible at 6.5 has a specific weakness you did not know about, and finding it before test day is the entire point.
Does this work if a different tool gave me the band?
Yes — paste the essay, the band, and the rubric into a fresh chat. It works slightly better across tools, because the second model has no earlier answer to protect.