System card
32 min
Kerrfield One system card
How Kerrfield One was trained, what 11,400 hours of red-teaming found, and every mitigation we shipped before launch.
Kerrfield safety team

This is the full record of how Kerrfield One was built and tested. We publish it on launch day because the people deciding whether to trust a model should be able to check our work.
Over eight weeks, external teams in biology, cybersecurity and persuasion spent 11,400 hours trying to make the model cause harm. They found 212 distinct failures. We fixed 197 before release and describe the other 15, with the reasons we shipped anyway.
What’s inside
Training data sources and filtering, by stage
Every evaluation, with prompts and raw outputs
Red-team findings grouped by severity
Known failure modes we have not fixed
Read it with a sceptical eye. If you find something we missed, the bug bounty pays up to $250,000.
Keep reading
More from the lab.

System card
32 min
Kerrfield One system card
How Kerrfield One was trained, what 11,400 hours of red-teaming found, and every mitigation we shipped before launch.

Paper
18 min
Thinking longer, not bigger
A 70B model given four minutes to think beats a model ten times its size answering at once. Here is why, and where it stops working.

Paper
14 min
Calibration you can check
When Kerrfield One says it is 80% sure, it is right about 80% of the time. How we trained for honest confidence, and how to verify it yourself.