Generate True/False/Not Given items with a named trap each
Forces the model to choose one of seven trap types before writing each statement, then prove the answer with the deciding sentence from the passage.
Using ONLY this passage: [PASSAGE] Write 8 True/False/Not Given statements. For each, you must assign one trap type BEFORE writing the statement, and the statement must actually implement it: - T1 Paraphrase-only (True, heavy paraphrase, no trick) - T2 Overstated (passage hedges, statement asserts → False) - T3 Reversed cause (statement swaps cause and effect → False) - T4 Scope creep (passage says "some", statement says "all" → False) - T5 Plausible-but-absent (world-knowledge true, passage silent → Not Given) - T6 Partial match (half in the passage, half not → Not Given) - T7 Comparison not made (passage describes both, never compares → Not Given) Output the 8 statements alone. Then a line of dashes. Then the key: answer, trap type, and the exact sentence(s) from the passage that decide it. Self-check before answering: for every NOT GIVEN, confirm the passage genuinely does not settle the claim. If it does, rewrite the item.
Best on Claude · ChatGPT · Gemini.
What to fill in
| Token | What to paste |
|---|---|
| [PASSAGE] | A single academic passage of 700–900 words. Do not paste several — the Not Given logic breaks when the model can borrow facts from elsewhere. |
What you should get back
Eight statements with no answers visible, then a key giving the answer, the trap type it implements, and the exact passage sentence that decides it.
What it actually returned
Tested on Gemini 2.5 Flash, API, August 2026, 2026-08-06. Trimmed — this is an excerpt of a longer response, not the whole thing.
1. The domestication of horses had a more significant impact on human civilization than nearly all other animal partnerships. 2. The wear patterns found on horse teeth from the Botai culture definitively prove that horse riding began around 3500 BCE. 3. Changes in land use patterns led to the development of mounted herding practices. 4. All societies across Eurasia quickly adopted horses as their primary mode of transport once they became available. 5. Horses were primarily used for warfare before they were used for agriculture. ——— 1. True, T1 Paraphrase-only, "The domestication of the horse transformed human society more profoundly than almost any other animal partnership."
Our verdict on this run
The trap discipline worked. Item 2 hardens the passage's "some researchers argue" into "definitively prove" (overstated → False), item 3 reverses the causal direction the passage gives — mounted herding altered land use, not the other way round — and item 4 turns "populations across the Eurasian steppe" into "all societies across Eurasia". Those are three distinct, correctly implemented trap types, which is more than most AI-generated True/False/Not Given manages. It also held the key back below a divider as instructed. Check item 5 yourself: the passage says the chariot preceded cavalry but never compares warfare with agriculture, which makes the Not Given defensible but the reasoning worth verifying.
How this prompt goes wrong
Invalid Not Givens. This is the single most common failure in AI-generated reading practice: the model writes a statement it labels Not Given when the passage actually contradicts it (that is False), or labels something False that the passage never addresses. The self-check line reduces it but does not eliminate it. Use the key's "deciding sentence" as your audit — if the key cannot quote a sentence for a False, the item is probably a Not Given, and if it quotes one for a Not Given, the item is broken. Trust your own reading over the key when the two disagree and you can point at the sentence.
Tips
- The trap taxonomy is worth more than the drill. Learning to name why an item is False is what transfers to the exam.
- Cross-check every Not Given you disagree with. If you can argue it from the passage, the item is broken — that is not you being wrong.
- Do the eight statements under a four-minute timer, which is roughly real exam pace for this type.
Related prompts
- ReadingAdjudicate one item you got wrong, and find the misreadingTakes a single wrong answer and returns the deciding sentence, the False-versus-Not-Given logic, the exact feature you misread, and a 10-second in-exam test.
- ReadingBuild a paraphrase drill where no content word repeatsTurns any passage into 12 paraphrase-matching items with three deliberate near-misses, and a hidden key that names each distortion.
- Writing Task 2Force the model to quote your essay before it scores itMarks your Task 2 essay in four locked steps — evidence first, descriptors second, band last — so it cannot work backwards from a flattering number.
- Mock examinerA 250-word verdict with no room for flatteryA tired, strict examiner persona capped at 250 words: the band first, three quoted justifications, one upgrade condition, nothing else.
Common questions
Why is AI-generated True/False/Not Given so often wrong?
Because the False/Not Given boundary is a logic judgement about what a text does and does not settle, and models are trained to be helpful rather than strict. Naming the trap type before writing the statement, and demanding the deciding sentence, is what makes the items auditable.
What if I disagree with the key?
Check whether the key can quote a sentence that settles the claim. If it cannot, the item is a Not Given regardless of what the key says. Being able to make that argument is the skill the real exam tests.
Which trap type catches most candidates?
T5 — plausible but absent. The statement is true in the world, so it feels right, but the passage never says it. Every one of those you get wrong is a habit of answering from knowledge rather than from the text.