Sobre este puesto de RL Environments Engineer en Bespoke Labs
About Bespoke Labs
Bespoke Labs is an applied AI research lab pioneering data and RL environment curation for training and evaluating agents.
Recently, we curated Open Thoughts, one of the best open reasoning datasets used by multiple frontier labs, trained SOTA specialized models such as Bespoke-MiniChart-7B and Bespoke-MiniCheck, and taught agents to do multi-turn tool-calling with reinforcement learning.
Bespoke is uniquely positioned to capture a large market share of data and RL environment curation.
About the Role
You will own coding environments end to end. You choose what world to build, design the tasks inside it, build the grading, run frontier models against it, and harden it until the only way to pass is to actually do the work.
This is a delivery role. We care most about whether you have done this before and can point to what came out of it. If you have built agentic coding environments or tasks at real volume and can tell us how many and how hard they were, we want to talk.
What You'll Do
Build high-fidelity coding worlds around real codebases, with the conventions, dependencies, tooling, and accumulated mess that real software has.
Choose which environments are worth building. A strong coding environment hits several marks:
Targets work where frontier models measurably struggle
Exercises real engineering, meaning navigation, diagnosis, sequencing, and design, and not just writing a function
Rests on a codebase with enough history and structure that shortcuts do not survive
Has a clear pass condition that a reviewer would agree with
Produces many varied tasks from a single world rather than one
Design tasks across the full lifecycle. Prompt, environment, grader, running frontier models, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.
Build grading and sandboxed execution that is deterministic and cannot be gamed. Assume the model will try to pass without doing the work, and close the door before it finds it.
Remove whatever is slowing the team down. Build the internal tooling that makes everyone around you faster.
Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.
What We're Looking For
A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.
Experience scaling that output through automation rather than through more people doing more manual work.
Strong software engineering fundamentals and fluency in several languages that holds up in production code.
Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.
An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.
A clear sense of what frontier coding agents can and cannot do, and where they cut corners.
Ownership. You build, debug, and ship without much supervision.
You May Be a Good Fit If You Also
Have worked on RL training systems, post-training, verifiers, or tool-use harnesses
Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure
Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment
Have contributed to a public agentic benchmark such as Terminal-Bench
Have open-source work that other people depend on
What We Offer
Location: Mountain View, CA (Onsite)
Base Salary: $250,000 – $300,000 USD / year
Additional Comp: 25% performance-based bonus + equity
Benefits & Perks
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
Visa sponsorship and relocation support available
Direct impact on how the industry trains and evaluates agents
We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.