Paper

18 min

Thinking longer, not bigger

A 70B model given four minutes to think beats a model ten times its size answering at once. Here is why, and where it stops working.

M. Castell, J. Okafor, T. Lindqvist

The usual way to make a model smarter is to make it bigger. We show that for hard problems, giving a model more time to search, test and revise is a cheaper lever.

On our hardest maths and science set, accuracy rises from 41% at one second to 92% at thirty minutes. A model ten times larger, answering immediately, reaches 63%.

What we found

  • Gains come mostly from dropping bad branches early

  • Self-checking matters more than longer chains

  • Returns flatten on tasks with no way to verify an answer

The code and evaluation harness are open. Everything needed to reproduce the curves in this paper is in the repository.

Start here

Ask it somethinghard.

Kerrfield
An independent research lab building models that reason before they answer.
© 2026 Kerrfield Research Ltd. · London · MontréalAll systems normal

Create a free website with Framer, the website builder loved by startups, designers and agencies.