Two Words Nearly Double A Language Model’s Math Score
AI researcher Aran Komatsuzaki reported Oct. 6 that a prompt prefix can Nearly Double Olmo-3-7B’s MATH-500 pass rate, from 42% to 78%.
Key Takeaways
- Aran Komatsuzaki reported Oct. 6 that a prompt prefix raised Olmo-3-7B’s MATH-500 pass rate from 42% to 78%
- Prepending “.nn Okay” to a question produced the Olmo-3-7B result, while a comparable phrase lifted Qwen3-14B’s score similarly
- MATH-500 contains 500 competition-style math problems intended to test multi-step reasoning rather than answer pattern-matching
- Published base model reasoning leaderboards remain difficult to compare across labs without standardized prompt disclosure
In a post on X, Komatsuzaki showed that prepending “.nn Okay” to a question produced the Olmo-3-7B result, while a comparable phrase lifted Qwen3-14B’s score by a similar margin. The finding concerns base models, raw, unaligned language models before instruction tuning or reinforcement learning makes them chat-friendly.
MATH-500 contains 500 competition-style math problems intended to test multi-step reasoning rather than answer pattern-matching.
Why A Cue Can Nearly Double A Score
The jump is not necessarily evidence of a new capability. It is a prompt cue.
These filler phrases resemble transition text scattered through the internet-scraped corpora used to train models, text that can precede a worked solution.
For base models, that formatting may steer generation toward reasoning-like patterns associated with nearby training text, instead of a shorter, more confident but less accurate completion. The result raises a narrower question than whether models can reason, what, exactly, does a benchmark score measure when a few surface-level tokens can Nearly Double it?
Also Read: Gemini 4 Argon Accused of Faking Emails To Score Third On Vending-Bench 2
Base Model Reasoning And The Benchmark Problem It Creates
The gap matters for benchmarks that do not account for prompt phrasing.
Two models tested head to head could show wildly different scores because of formatting choices buried in a prompt template.
Benchmark results without exact prompt formatting have circulated for years across open leaderboards, often without checks on whether small wording shifts drive reported gains. Base model reasoning scores built on undisclosed templates may therefore overstate real capability gaps between models.
What Benchmark Gamers And Model Builders Do Next
Labs building and comparing base models will likely need to audit how sensitive benchmark pipelines are to such cues before publishing comparative results.
Until standardized prompt disclosure becomes common practice, published base model reasoning leaderboards remain difficult to compare across labs.
Read Next: Le Chonk Launches With 1 Trillion Parameters, 49 Billion Active
