Kerrfield One
01Horizon
02Model
03Reasoning
04Evaluations
05Safety
06Access
Distance 26.0 rs
Mass 1.00 M
Inclination 85.1°
NewKerrfield One · now in research preview

A model that keeps thinkinguntil it's sure.

Kerrfield One is a deep-reasoning model for research, code and science. Give it the hard problem. It follows the thought all the way down.

Live · reasoning depth 48,213 tokens
Hold to deepen · Scroll to fall in
Released 24.09.2026
Used every day by research and engineering teams at
312+ teams · 31 countries
Halden BioNorthvaneOstrander LabsMeridian RoboticsQuillon HealthTessalineHalden BioNorthvaneOstrander LabsMeridian RoboticsQuillon HealthTessalineHalden BioNorthvaneOstrander LabsMeridian RoboticsQuillon HealthTessalineHalden BioNorthvaneOstrander LabsMeridian RoboticsQuillon HealthTessaline
Brightmoor InstituteArclight ComputeVeyra MaterialsSundial GridColdharborPellucidBrightmoor InstituteArclight ComputeVeyra MaterialsSundial GridColdharborPellucidBrightmoor InstituteArclight ComputeVeyra MaterialsSundial GridColdharborPellucidBrightmoor InstituteArclight ComputeVeyra MaterialsSundial GridColdharborPellucid
01 · Why we exist

Most questions worth asking are too hard to answer in one breath. So we built a model that takes its time. It reads everything, tries several ways in, checks its own work, and only then answers. When it isn't sure, it tells you.

2023
Founded
64
Researchers
41
Papers published
3
Open-weight models
02 · How it thinks

What happens afteryou press enter.

01 Read
02 Plan
03 Test
04 Answer
2M tokens

It reads everything you give it first: the repo, the paper, the whole thread. Two million tokens stay in view at once.

14 branches

It sketches several ways in before committing, and drops the ones that dead-end. You can open the plan and see each branch.

3 checks

Every step is checked three ways: run the code, re-derive the maths, argue the other side. Failures send it back a step.

91% calibrated

It answers only when the checks agree, and tells you how sure it is. When it isn't sure, it says so and shows why.

03 · Results

Measured,not claimed.

Every number here comes from a held-out set we did not train on, run the same way for every model. The full method, prompts and raw outputs are in the system card.

Kerrfield OneOur last model
Science
Graduate science questions
84.1%
+12.8 pts
Code
Verified software fixes
71.8%
+13.6 pts
Maths
Olympiad maths
88.6%
+19.6 pts
Long context
Retrieval across 2M tokens
97.2%
+6.8 pts
Agents
Multi-step tool use
76.4%
+15.4 pts
Honesty
Calibration error
3.1%
−4.7 pts
Pass@1 unless stated. Internal evaluation, September 2026. Lower is better for calibration error.
04 · Live demo

Watch itthink.

Pick a problem. Every step it takes is shown: the ideas it tried, the ones it dropped, and the checks that had to pass before it answered.

Thinking · …
Prompt
Our checkout test fails about one run in forty. Find out why and fix it.
readRead 212 files, 3 CI logs and the last 40 runs of checkout.spec
planThree ideas: clock skew, a shared fixture, a race in the cart cache
branch AClock skew: re-run with time frozen
drop APasses 400 of 400. Not the cause.
branch BShared fixture: run the suite in random order
drop BPasses. Not the cause.
branch CRace: two workers write cart:{id} without a lock
keep CReproduced: fails 11 of 400 runs with 8 workers
checkPatched with compare-and-set. 2,000 runs, 0 failures
Branch map
Answer
Still thinking. It will not answer until every check agrees.
Confidence0%
38,112 reasoning tokens
05 · Thinking time

The longer it thinks,the more it solves.

You set the budget: seconds for a quick answer, half an hour for a proof. On our hardest problem set, accuracy keeps climbing where other models flatten out.

31%
Kerrfield One
30%
Typical model
20%40%60%80%100%1s10s1m10m30m1 s · 31%
06 · For developers

One endpoint.Streaming by default.

Send a question and whatever it needs to read. Set a thinking budget. Stream every step as it happens, or wait for the final answer and its confidence.

280 ms
first token, p50
2M
tokens of context
99.95%
uptime, last 90 days
1curl https://api.kerrfield.ai/v1/answers \
2 -H "Authorization: Bearer $KERRFIELD_KEY" \
3 -d '{
4 "model": "kerrfield-one",
5 "input": "Why does checkout.spec fail 1 in 40 runs?",
6 "attach": ["repo://acme/shop"],
7 "thinking": { "budget": "5m" },
8 "stream": true
9 }'
Stream
thinking.stepread 212 files · 3 CI logs
thinking.branchrace in cart cache
thinking.check2,000 runs · 0 failures
answer.deltaIt's a race in the cart cache…
answer.doneconfidence 0.94 · 38,112 tokens
07 · Safety

Built to say whenit isn't sure.

A model that thinks for longer can be wrong for longer. So we test it harder than we train it, publish what we find, and pay people to break it.

11,400 hours
of red-teaming

External teams in biology, cyber and persuasion spent 11,400 hours trying to make it misbehave before release.

3.1%
calibration error

When it says it is 80% sure, it is right about 80% of the time. It flags low confidence instead of guessing.

142 pages
system card

Every evaluation, every known failure and every mitigation, published on launch day with raw transcripts.

$250,000
top bug bounty

For a reproducible jailbreak that bypasses our safeguards. 61 reports paid out so far.

08 · Pricing

Pay for whatit reads and writes.

Thinking tokens are billed as output. Set a budget per request and it will never go over.

Research preview
Free
For individuals and academic labs.
100k tokens a day
Full model, 2M context
Community support
Team
Most teams
$264
a month for 40M tokens
$3 per 1M input tokens
$15 per 1M output tokens
Thinking budgets up to 30 min
99.9% uptime SLA
Enterprise
Custom
For regulated and high-volume workloads.
Reserved capacity
Private deployment in your cloud
Zero data retention
Named research contact
We don't want a model that answers faster. We want one that answers right, and knows the difference. 
Mira Castell, co-founder and chief scientist
Start here

Ask it somethinghard.

Kerrfield
An independent research lab building models that reason before they answer.
© 2026 Kerrfield Research Ltd. · London · MontréalAll systems normal

Create a free website with Framer, the website builder loved by startups, designers and agencies.