Turn a recording of yourself into countable speaking data
Counts your words per minute, filled pauses, restarts, repeated words and structure range from a transcript — and refuses to guess a band.
This is a transcript of me speaking in an IELTS mock test: [TRANSCRIPT] Official Speaking band descriptors: [RUBRIC] Analyse only what is measurable in the transcript. Give me: 1. Fluency markers: total words, words per minute, count of filled pauses (um/er), count of self-corrections, count of repeated restarts, longest unbroken stretch. 2. Lexical resource: count of distinct content words, my 10 most repeated content words, and 5 places where a more precise word existed (quote me, then suggest). 3. Grammatical range: list every structure type I used (present simple, present perfect, conditional, relative clause, passive...) with counts. Name the three structures a Band 7+ speaker would use here that I never used once. 4. Coherence: quote every discourse marker I used. Flag over-use. 5. Then: the three habits costing me the most, in order. Do not estimate a band. Do not comment on pronunciation — you cannot hear it.
Best on ChatGPT · Claude · Gemini.
What to fill in
| Token | What to paste |
|---|---|
| [TRANSCRIPT] | A transcription of your own recorded answer, including the "um"s. Any free speech-to-text tool will do; do not clean it up. |
| [RUBRIC] | The official public Speaking band descriptors from ielts.org. |
What you should get back
A quantified profile of your own speech — counts, your ten most repeated content words, the structures you never used once — and three ranked habits to fix. No band score.
What it actually returned
Tested on Gemini 2.5 Flash, API, August 2026, 2026-08-06. Trimmed — this is an excerpt of a longer response, not the whole thing.
1. Fluency Markers:
— Total words: 143 words
— Words per minute (WPM):
Assuming a 1 minute 30 second response: 95.3 WPM
Assuming a 2 minute response: 71.5 WPM
(Please note: WPM is an estimation as the exact duration of your speech is not provided in the transcript.)
— Count of filled pauses (um/er): 5 ("Um", "um", "um", "um", "um")
— Count of self-corrections: 1 ("it's called... it was")
— Count of repeated restarts: 3 ("and, and their father", "they, they discover", "I can, I can relate")
— Longest unbroken stretch: "I liked it because the writing is very simple but it makes you feel a lot of things." (20 words)Our verdict on this run
It refused to invent the one number it could not know. Rather than assert a words-per-minute figure, it gave the arithmetic for two possible durations and said why — which is better behaviour than most marking prompts produce, and a reminder to record your duration alongside the transcript. The restart and self-correction quotes are accurate and genuinely trackable month to month. It also obeyed the two prohibitions: no band, and no comment on pronunciation. Treat the counts as close rather than exact; models are not reliable counters and the totals shift slightly between runs.
How this prompt goes wrong
The counts are approximate. Language models do not count reliably, so words per minute and filler totals can be out by 10–20%, and asking twice can give two answers. That does not ruin the prompt — the value is in the ratios and the trend across weeks, not the absolute number — but do not quote these figures as if they were measured. The bigger failure is the model ignoring the last line and offering a band anyway. Discard it: a transcript carries no pronunciation, which is a quarter of the Speaking mark, so any band derived from text alone is fabricated.
Tips
- Record the same cue card monthly and compare the numbers. The trend is trustworthy even when the individual counts are not.
- The "three structures a Band 7+ speaker would use that you never used once" is the highest-value line — take one and drill it deliberately.
- Do not clean up the transcript before pasting. The false starts are the data.
Related prompts
- SpeakingSee the same cue card answered at Band 6, 7 and 8Produces three spoken-register answers to one cue card at ascending bands, then names the concrete moves that separate each level from the next.
- SpeakingMake the examiner push back on every Part 3 answerSix escalating Part 3 questions with real counter-examples and follow-ups, then a report on where your answers were too short or unsupported.
- Mock examinerRun a full Speaking test where the examiner never breaks roleA complete Parts 1–3 interview in examiner register — no teaching, no encouragement, no band — so it feels like the real thing.
- Writing Task 2Force the model to quote your essay before it scores itMarks your Task 2 essay in four locked steps — evidence first, descriptors second, band last — so it cannot work backwards from a flattering number.
Common questions
Why will it not give me a band?
Because the prompt forbids it, deliberately. Pronunciation is one of the four Speaking criteria and a transcript contains none of it, so a band from text is a guess dressed up as a measurement. Use a tool that hears your audio if you want a band.
How do I transcribe my answer?
Any free speech-to-text tool works — your phone's voice typing, a browser dictation feature, or the transcription built into most note apps. Accuracy matters less than you would think, because the fluency markers survive small errors.
What words-per-minute should I aim for?
There is no target in the descriptors. Natural IELTS speech usually lands somewhere between 120 and 160, but a slower speaker who develops ideas fully scores above a fast one who does not. Track your own trend rather than chasing a number.