The model · Kerrfield One

A model thatchecks its own work.

Kerrfield One reads everything you give it, tries several ways in, tests each one, and answers only when the checks agree. Use it in the app or on the API.

2M tokens
Context
1 s – 30 min
Thinking
3.1%
Calibration error
24.09.26
Released
Capabilities

What it isgood at.

01
Long, messy inputs

Give it the whole codebase, the full paper trail or a year of logs. It reads all of it before answering.

02
Code that has to work

It runs what it writes in a sandbox, reads the errors and tries again before handing it over.

03
Maths and science

Olympiad-level problem solving with every step shown, so a reviewer can check the working.

04
Tools and agents

Browses, queries databases and calls your APIs, with a plan you can read before it acts.

05
Honest uncertainty

Every answer carries a confidence score. Below your threshold, it asks instead of guessing.

06
Private by default

API data is never used for training. Zero-retention and in-VPC deployment on request.

Specification

Kerrfield One,in numbers.

Everything a developer needs before the first call. Values are for the hosted API; private deployments can raise limits.

Context window
2,000,000 tokens
about 1,500 pages or a mid-sized repository
Max output
128,000 tokens
including visible reasoning
Thinking budget
1 s to 30 min
set per request, never exceeded
Modalities
Text, code, images, PDFs
audio in preview
Knowledge cut-off
June 2026
plus live web and file tools
Latency
280 ms first token
p50, 5 s budget
Languages
41
evaluated to within 5% of English
Weights
Closed · open 8B distill
the distill is Apache 2.0
Results

Measured,not claimed.

Every number here comes from a held-out set we did not train on, run the same way for every model. The full method, prompts and raw outputs are in the system card.

Kerrfield OneOur last model
Science
Graduate science questions
84.1%
+12.8 pts
Code
Verified software fixes
71.8%
+13.6 pts
Maths
Olympiad maths
88.6%
+19.6 pts
Long context
Retrieval across 2M tokens
97.2%
+6.8 pts
Agents
Multi-step tool use
76.4%
+15.4 pts
Honesty
Calibration error
3.1%
−4.7 pts
Pass@1 unless stated. Internal evaluation, September 2026. Lower is better for calibration error.
Thinking time

The longer it thinks,the more it solves.

You set the budget: seconds for a quick answer, half an hour for a proof. On our hardest problem set, accuracy keeps climbing where other models flatten out.

31%
Kerrfield One
30%
Typical model
20%40%60%80%100%1s10s1m10m30m1 s · 31%
System card

142 pages. Every test,every known failure.

How Kerrfield One was trained, what we tested, what went wrong and what we changed. Includes raw red-team transcripts and our evaluation harness.

Download the system cardkerrfield-one-system-card.pdf · 8.4 MB
Questions

Asked often,answered plainly.

No. Nothing sent through the API or the app is used to train our models. Enterprise plans add zero data retention.

Thinking tokens are billed as output tokens. You set a budget per request, and a request never spends more than its budget.

Yes, on Enterprise. We deploy into your AWS, Azure or GCP account and never see your traffic.

It says so. Every answer carries a confidence score, and you can set a threshold below which it asks a clarifying question instead.

Yes. Kerrfield One Distill (8B) is released under Apache 2.0 for research and local use.

London and Montréal. Our research team is spread across 11 time zones.

Start here

Ask it somethinghard.

Kerrfield
An independent research lab building models that reason before they answer.
© 2026 Kerrfield Research Ltd. · London · MontréalAll systems normal

Create a free website with Framer, the website builder loved by startups, designers and agencies.