Skip to content

What the “Repeat the Prompt Twice” Study Actually Found

A bounded reading of the 2025 prompt-repetition paper, including non-reasoning conditions, statistical comparisons, reasoning results, and long-input exceptions.

Published: Updated: Reviewed: Author: Category: AI for work and thinking

2 min read

Checking the conditions of the prompt-repetition study

“Write every question twice and every AI becomes more accurate” is not supported. The December 2025 paper Prompt Repetition Improves Non-Reasoning LLMs compared specific models, benchmarks, and prompt arrangements, mainly when models were told not to reason.

What was compared

The transformation changes <QUERY> to <QUERY><QUERY>. The paper evaluated seven models and seven benchmarks in multiple configurations. Models included Gemini 2.0 Flash variants, GPT-4o variants, Claude 3 variants, and DeepSeek V3; API tests ran in February and March 2025.

ConditionReported resultCorrect interpretation
Reasoning disabled47 wins, 0 losses across 70 model-benchmark configurationsThe remaining comparisons had no statistical win; not every answer improved
Step-by-step reasoning encouraged5 wins, 1 loss, 22 neutral across 28 comparisonsThe effect was smaller and not universally positive
Long requestsHigher latency in some Claude conditionsRepetition still doubles input and can approach context limits

The win/loss labels use the paper's McNemar test and significance threshold. “47 wins” does not mean 47 real-world uses were guaranteed to succeed.

Why it may help

The authors focus on causal language models: earlier tokens cannot attend to later tokens. In the repeated copy, each part of the second query has the entire first query as earlier context. This is the proposed mechanism, not a correctness check for an individual answer.

Evaluate it on your own task

  1. Create at least 20 representative cases with a checkable answer.
  2. Hold the model, settings, temperature, and output format constant.
  3. Compare the original and repeated prompts on the same cases.
  4. Record accuracy, input tokens, latency, and format violations.
  5. Do not adopt repetition when it provides no useful gain.

Before production use, consider duplicated sensitive text and the larger input size.

Poor fits

  • Inputs already near a context limit.
  • Creative work without an objective success measure.
  • Reasoning settings that already reread or restate the problem.
  • Tool systems where duplicated text could be misunderstood as two requested actions.

Conclusion

Prompt repetition is a bounded input transformation that performed well in the paper's main non-reasoning tests. It is not a universal accuracy switch. Test quality, cost, latency, and format on your own cases before adopting it.

Primary sources checked

Important claims should also link to the relevant source in the article body.

  1. Prompt Repetition Improves Non-Reasoning LLMsarXiv · research-paper · Checked: 2026-07-26

Related posts

Author

ImidefWorks

An independent writer who connects primary sources with reproducible checks across AI, web publishing, development, and information organization.

View author profile and editorial policy