Skip to content
Back to Blog
Speaking Strategies

How to Practise IELTS Speaking With AI (And What It Cannot Hear)

JM

Jahidul Hossain Mekat

Head of AI & Computational Linguistics at IELTSbiz

August 6, 20268 min read

Key takeaways

  • A text model cannot hear you, and pronunciation is one of the four Speaking criteria — any band it gives you is missing 25% of the mark.
  • What AI does genuinely well: unlimited Part 3 pushback, and counting fluency markers from a transcript that you cannot count yourself.
  • The most useful numbers are words per minute, filled pauses, restarts, and the structures you never used once.
  • Track those numbers on the same cue card a month apart. Change is the measurement that matters, not level.
  • The failure mode to watch for is role collapse — the moment it says "great answer!", it stops resembling the exam.

AI is the best sparring partner most IELTS candidates will ever have access to, and it is deaf. It will run a Part 3 discussion at two in the morning, challenge every general claim you make, and never get bored of the topic.

It cannot hear that your intonation flattens when you are nervous, or that a consonant cluster is costing you intelligibility. Since pronunciation is one of the four Speaking criteria, that is a quarter of your mark it has no access to.

Knowing exactly where the line sits is what separates candidates who get real value out of AI speaking practice from those who get a false sense of readiness.

What AI genuinely does well

1. Pressure, on demand

Part 3 is where Band 6.5 becomes Band 7.5, and the deciding skill is extending and defending a position rather than deploying impressive vocabulary.

A model instructed to challenge every general claim with a counter-example, press you when you hedge, and never encourage you produces something genuinely close to the real experience. You can run it six times in an evening on six different themes.

The instruction that matters most is the negative one: no praise, no correction, no vocabulary help during the test. Without it you get a friendly conversation, which is the opposite of useful — the absence of reassurance is exactly what produces test-day nerves, and rehearsing without it rehearses the wrong thing.

2. Counting what you cannot count

Record yourself answering a cue card, transcribe it with any free speech-to-text tool, and hand the transcript to a model.

It will give you total words, filled pauses, self-corrections, restarts, your ten most repeated content words, and — the most useful line in the whole output — the grammatical structures a Band 7+ speaker would have used here that you never used once.

What a transcript can showWhat it cannot
Filled pauses, restarts, self-correctionsPronunciation of individual sounds
Vocabulary range and repetitionWord and sentence stress
Which grammatical structures you usedIntonation and its effect on meaning
Whether you answered the question askedWhether a listener actually understood you
How far you developed each answerAnything about your accent's intelligibility

The right-hand column is not a minor gap. It is a whole criterion, and it is the one candidates most often assume is fine.

The numbers worth tracking

Language models are unreliable counters — ask twice and the totals shift slightly — so treat these as approximations. Their value is in the trend, not the absolute figure. Record the same cue card today and again in four weeks, run both transcripts through the same analysis, and compare:

  • Filled pauses per hundred words. The single most visible fluency marker.
  • Restarts and self-corrections. A few are natural and even good; a cluster of them signals you are planning mid-sentence.
  • Longest unbroken stretch. Improves faster than almost anything else with practice.
  • Distinct content words. A crude but honest proxy for lexical range.
  • Structures never used. The gap list, and the best source of drills.

There is no target words-per-minute in the descriptors. Natural IELTS speech usually lands between roughly 120 and 160, but a slower speaker who develops ideas fully scores above a fast one who does not.

The failure mode to watch for

Every consumer assistant is tuned to be encouraging, and encouragement destroys this exercise. The tell is a "great answer!" between questions, or an explanation of what you should have said. The moment either appears, the model has left the examiner role and you are having a language lesson, not a mock test.

Restating the rule mid-conversation rarely fixes it. Put the role instructions in your tool's custom instructions or system prompt before the conversation starts, and begin again.

There is a second, subtler failure in text mode: the model writes the whole test as a script, stubbing your answers as "[User speaks]" and running through to the end without ever waiting for you. Voice mode avoids that entirely, because turn-taking is enforced by the medium.

We document both failures, with the real captured output, on our mock examiner prompt pages.

A routine that works

  1. Pick a card from our cue card library and answer it in voice mode with the full examiner prompt, recording as you go.
  2. Transcribe the recording and run the transcript analysis prompt. Save the numbers in a note.
  3. Run a Part 3 sparring session on the same theme, exactly as the real test does.
  4. For pronunciation and a band you can plan around, use a tool that processes your actual audio — our Speaking mock test does this on your recording rather than a transcript.
  5. Repeat the same card in four weeks and compare the two sets of numbers.

If you are preparing entirely on your own without AI in the loop, our guide to practising IELTS Speaking alone covers the offline version of the same routine.

JM

Jahidul Hossain Mekat

Head of AI & Computational Linguistics at IELTSbiz

LinkedIn Profile

Jahidul Hossain Mekat leads AI and computational linguistics at IELTSbiz, building the automated grading and feedback systems behind the writing checker and reading practice.

View all articles by Jahidul Hossain Mekat

Frequently Asked Questions

Can AI assess my IELTS Speaking?

It can assess what is in a transcript — fluency markers, vocabulary range, grammatical range and coherence — but not pronunciation, which is one of the four criteria. A band from a text-based model is missing a quarter of the mark. Tools that process your actual audio are a different matter and can give a more complete estimate.

Is practising IELTS Speaking with ChatGPT worth it?

For format familiarity, question exposure and Part 3 pushback, genuinely yes — it is repeatable at no cost and available whenever you are. For pronunciation feedback and the interactive judgement of a real examiner, it is not a substitute. Use it for volume and use something that hears you for the band.

How do I transcribe my speaking practice?

Your phone's voice typing, a browser dictation feature, or the transcription built into most note apps all work. Do not clean up the result — the fillers, restarts and self-corrections are exactly the data the analysis depends on.

Why does the AI keep praising me during a mock test?

Encouragement is the default behaviour of every consumer assistant, and a single instruction inside one message is weak against it. Put "no praise, no correction, stay in role" into your tool's custom instructions before the conversation starts — rules set up front hold far better than corrections issued mid-test.

Can AI give me an IELTS Speaking band score?

A text model will offer one if you ask, and you should not use it. Pronunciation is a quarter of the Speaking mark and a transcript carries none of it, so the number is partly invented. Prompts worth using refuse to give a band from text for exactly that reason.

Related posts

Ready to achieve your target IELTS score?

Practice with unlimited AI-generated Cambridge-style passages, receive instant examiner-level feedback, and track your band score progress.

    Need help?