Most questions worth asking are too hard to answer in one breath. So we built a model that takes its time. It reads everything, tries several ways in, checks its own work, and only then answers. When it isn't sure, it tells you.
Measured,not claimed.
Every number here comes from a held-out set we did not train on, run the same way for every model. The full method, prompts and raw outputs are in the system card.
Watch itthink.
Pick a problem. Every step it takes is shown: the ideas it tried, the ones it dropped, and the checks that had to pass before it answered.
The longer it thinks,the more it solves.
You set the budget: seconds for a quick answer, half an hour for a proof. On our hardest problem set, accuracy keeps climbing where other models flatten out.
One endpoint.Streaming by default.
Send a question and whatever it needs to read. Set a thinking budget. Stream every step as it happens, or wait for the final answer and its confidence.
Built to say whenit isn't sure.
A model that thinks for longer can be wrong for longer. So we test it harder than we train it, publish what we find, and pay people to break it.
External teams in biology, cyber and persuasion spent 11,400 hours trying to make it misbehave before release.
When it says it is 80% sure, it is right about 80% of the time. It flags low confidence instead of guessing.
Every evaluation, every known failure and every mitigation, published on launch day with raw transcripts.
For a reproducible jailbreak that bypasses our safeguards. 61 reports paid out so far.
Pay for whatit reads and writes.
Thinking tokens are billed as output. Set a budget per request and it will never go over.
09 · From the lab
Papers, notes and releases.

System card
32 min
Kerrfield One system card
How Kerrfield One was trained, what 11,400 hours of red-teaming found, and every mitigation we shipped before launch.

Paper
18 min
Thinking longer, not bigger
A 70B model given four minutes to think beats a model ten times its size answering at once. Here is why, and where it stops working.

Paper
14 min
Calibration you can check
When Kerrfield One says it is 80% sure, it is right about 80% of the time. How we trained for honest confidence, and how to verify it yourself.