The top AI labs train their models on data we build.
A small set of companies drives model progress from behind the scenes, building the RL environments frontier models train in. Diffuse Labs is one of them. We're a few people in San Francisco working directly with several of the top labs.
Research Engineer, RL Environments
Full-time · In person, San Francisco · Contract-to-hire is fineWe build reinforcement learning environments that frontier labs use to train their models. Most of our engineering work goes into generating tasks at scale, grading them reliably, and measuring whether training on them actually improves a model. As one of the first engineers, you will own large parts of this pipeline and work directly with our lab customers.
- Build agent pipelines that turn merged PRs, agent traces, and CI failures into containerized training environments with automatically checkable rewards.
- Design grading logic that holds up when a model is trained against it, and audit trajectories for reward hacking.
- Run rollouts across frontier models to measure task difficulty, and regenerate task sets as new releases saturate them.
- Run training ablations to quantify how much our environments improve a model on held-out evals.
- Build task sets for new domains as lab needs change. Recent examples include chip design, Kubernetes operations, and enterprise software.
- You use coding agents and frontier models daily and have good intuitions about where they fail.
- You are comfortable across the stack: containers, CI, data pipelines, and evaluation harnesses.
- You can read an agent trajectory and notice when a model is gaming its grader.
- RL experience is helpful but not required.
- A 30-minute call with a founder.
- A paid working session on a task from our production pipeline. Using AI tools is expected. We share results the same day.
- A short paid work trial, typically one to two weeks.
- Offer.
Different specialty? Write to team@diffuselabs.ai.
Life at Diffuse Labs
We're a few people working in person in San Francisco. Everyone owns a whole problem, talks to the labs directly, and ships work that ends up in frontier training runs. It's early enough that the people who join now decide what this company becomes.