What it isgood at.
Give it the whole codebase, the full paper trail or a year of logs. It reads all of it before answering.
It runs what it writes in a sandbox, reads the errors and tries again before handing it over.
Olympiad-level problem solving with every step shown, so a reviewer can check the working.
Browses, queries databases and calls your APIs, with a plan you can read before it acts.
Every answer carries a confidence score. Below your threshold, it asks instead of guessing.
API data is never used for training. Zero-retention and in-VPC deployment on request.
Kerrfield One,in numbers.
Everything a developer needs before the first call. Values are for the hosted API; private deployments can raise limits.
Measured,not claimed.
Every number here comes from a held-out set we did not train on, run the same way for every model. The full method, prompts and raw outputs are in the system card.
The longer it thinks,the more it solves.
You set the budget: seconds for a quick answer, half an hour for a proof. On our hardest problem set, accuracy keeps climbing where other models flatten out.