Skip to content

IELTS Mock examiner AI prompts

Personas that stay in role, stay strict, and stay quiet.

Every consumer chatbot is tuned to be encouraging, and encouragement is the single biggest problem with using one as an examiner. It praises before it assesses, it softens the verdict, it breaks character to teach in the middle of a test, and each of those things makes the exercise less like the exam rather than more. The prompts in this category are mostly constraints, because constraints are what hold a persona in place.

Three mechanisms do most of the work. A word limit — flattery needs room to exist, and a 250-word cap forces the model to pick the three things that actually decide the band. An explicit silence instruction, so a timed writing simulation gives you the question and then says nothing at all, which is what makes it a test rather than a lesson. And a role rule set before the conversation starts rather than during it, because instructions issued mid-conversation rarely stick once a model has settled into a helpful register.

Framing moves the number, and that is worth understanding before you trust any AI band. A model told it is a tired, difficult examiner will often return half a band lower than the same model given a neutral prompt on the same essay — not because it read the essay differently, but because it is playing a character. That gap is a measurement of how much the prompt, rather than your writing, is driving the score. Run both and look at the difference. The quoted justifications stay stable across framings; the digit does not.

For comparison: when we measured purpose-built AI grading against three certified and former IELTS examiners across 1,200 real Task 2 essays, bands agreed within ±0.5 in 94.2% of cases. That study is published with its method. Personas are a useful way to strip flattery out of a chatbot, but they are not a calibration, and the difference between the two matters when you are deciding whether you are ready to book the test.

The 4 prompts

Common questions

Does telling the AI to be a strict examiner make it more accurate?

It makes it less inflated, which is usually closer to the truth for a general chatbot, but it can overshoot — "strict" is a character instruction, not a calibration. Compare a strict run against a neutral one on the same essay; the gap tells you how much the framing is doing.

Why does the AI break character during a mock test?

Because helpfulness outranks the role instruction in its default behaviour. Setting the rules in custom instructions or a system prompt before the conversation begins holds far better than putting them in the first message.

Can an AI mock test replace a real one?

For format familiarity, timing and repetition, it is genuinely useful and free. For pronunciation, interactive judgement and a band you can plan around, it is not a substitute for a properly marked mock.

Continue

    Need help?