Inference engineer
Make long thinking cheap and fast: batching, caching and kernels for 30-minute requests.
Apply for this role
€140k–€200k + equity
Our requests can run for half an hour. Serving them well is an unsolved problem, and it is yours.
What you’ll do
Write and tune GPU kernels
Design scheduling for very long requests
Cut first-token latency further
You have shipped high-performance systems and like measuring before optimising.
Open roles