What the “Repeat the Prompt Twice” Study Actually Found
A bounded reading of the 2025 prompt-repetition paper, including non-reasoning conditions, statistical comparisons, reasoning results, and long-input exceptions.
2 min read

“Write every question twice and every AI becomes more accurate” is not supported. The December 2025 paper Prompt Repetition Improves Non-Reasoning LLMs compared specific models, benchmarks, and prompt arrangements, mainly when models were told not to reason.
What was compared
The transformation changes <QUERY> to <QUERY><QUERY>. The paper evaluated seven models and seven benchmarks in multiple configurations. Models included Gemini 2.0 Flash variants, GPT-4o variants, Claude 3 variants, and DeepSeek V3; API tests ran in February and March 2025.
| Condition | Reported result | Correct interpretation |
|---|---|---|
| Reasoning disabled | 47 wins, 0 losses across 70 model-benchmark configurations | The remaining comparisons had no statistical win; not every answer improved |
| Step-by-step reasoning encouraged | 5 wins, 1 loss, 22 neutral across 28 comparisons | The effect was smaller and not universally positive |
| Long requests | Higher latency in some Claude conditions | Repetition still doubles input and can approach context limits |
The win/loss labels use the paper's McNemar test and significance threshold. “47 wins” does not mean 47 real-world uses were guaranteed to succeed.
Why it may help
The authors focus on causal language models: earlier tokens cannot attend to later tokens. In the repeated copy, each part of the second query has the entire first query as earlier context. This is the proposed mechanism, not a correctness check for an individual answer.
Evaluate it on your own task
- Create at least 20 representative cases with a checkable answer.
- Hold the model, settings, temperature, and output format constant.
- Compare the original and repeated prompts on the same cases.
- Record accuracy, input tokens, latency, and format violations.
- Do not adopt repetition when it provides no useful gain.
Before production use, consider duplicated sensitive text and the larger input size.
Poor fits
- Inputs already near a context limit.
- Creative work without an objective success measure.
- Reasoning settings that already reread or restate the problem.
- Tool systems where duplicated text could be misunderstood as two requested actions.
Conclusion
Prompt repetition is a bounded input transformation that performed well in the paper's main non-reasoning tests. It is not a universal accuracy switch. Test quality, cost, latency, and format on your own cases before adopting it.
Primary sources checked
Important claims should also link to the relevant source in the article body.
- Prompt Repetition Improves Non-Reasoning LLMsarXiv · research-paper · Checked: 2026-07-26