Jobs Companies Nuro Applied AI Researcher, Agent Systems & Evaluation

À propos de ce poste Applied AI Researcher, Agent Systems & Evaluation chez Nuro

Nuro · Sur site · Mountain View, California (HQ)

Who We Are 

Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware. Nuro licenses its core technology, the Nuro Driver™, to support a wide range of applications, from robotaxis and commercial fleets to personally owned vehicles. With technology proven over years of self-driving deployments, Nuro gives automakers and mobility platforms a clear path to AVs at commercial scale, empowering a safer, richer, and more connected future.

About the Team

Frontier models are fungible. Any team can rent the same intelligence we can, and the model we build on today will be replaced within a month. What is not fungible is the infrastructure that decides whether an autonomous system's output can be trusted — evaluation, verification, and the discipline to gate on evidence instead of impressions. Nuro has spent a decade building exactly that discipline for a robot that drives on public roads, and this team turns it inward: we build the platform that lets AI agents operate autonomously inside Nuro's own engineering organization, under the same standard of proof we apply to the vehicle.

Our mandate is to amplify the output of every engineer and researcher at Nuro by 100x. Not a better IDE, not a faster build — a change in what a single person can attempt. That number is a target, not a claim, and reaching it depends on one thing above all: autonomous work has to be trustworthy enough to run unattended. So our central ambition is to build the most rigorous closed-loop evaluation system for AI work anywhere. Leverage follows from trust, and trust follows from measurement.

We operate as a startup inside a company that has already shipped a hard thing. Small team, no established playbook, direct access to compute and to the systems we are automating. You will work directly with engineering leadership and the CEO, and the decisions you make will be yours to make rather than yours to implement.

About the Role

Most teams building agents make design decisions by intuition and anecdote. Someone tries a new memory scheme, it feels better, it ships. We think that is the central failure of the field right now, and we are building this team to work the other way: every decision about how our agent systems are constructed should be settled by evidence.

You would not be starting from zero. We already operate a substantial agent system in production, a fleet of agents with an extensive library of skills and plugins, integrated into the tools our engineers use daily, serving real users with real work. So every hypothesis you form can be tested against genuine production traffic from your first month. And the system is now complex enough that intuition has stopped being sufficient to improve it, which is precisely why this role exists.

Everything this team builds is centered on frontier-lab transformer models. We are not inventing architectures. We are extracting the maximum from the best models that exist, and adapting them ourselves in the narrow places where our data gives us an advantage nobody else has.

Your charter has two halves. Make the system perform against real-world data, not public benchmarks or tasks we invented to look good, but our codebase, our infrastructure, and our engineers' actual requests, with all the ambiguity that implies. The gap between benchmark performance and real-task performance is where most agent systems quietly fail. And turn any task into a closed loop: for any workflow an agent takes on, you should be able to say what success looks like, where the evaluation data comes from, how signal is collected, and how results feed the next iteration.

Own the evaluation pipeline end to end. Three stages, and the value is in owning all three:

  • Eval data collection. Where ground truth comes from. Mining production traces for labeled outcomes, capturing human accept/reject/edit signal as it happens, building task sets that reflect the real distribution of work rather than the tasks that are easy to score, and knowing when a model-based judge is trustworthy and when it is laundering an assumption.
  • Eval loop construction. Turning a fuzzy objective into a measurement that runs on every change. Noise floors, statistical standards for acceptance, task suites that resist gaming, and experiments designed for production settings where clean randomization is not always available.
  • Automated hill climbing. The payoff. Once a task has a trustworthy loop, improvement can be searched rather than hand-crafted — prompts, context strategies, tool sets, routing, reasoning budgets, eventually model choice. This only works if the first two stages are sound; done wrong, it optimizes hard against a metric that means nothing. Post-train models on data nobody else has. This role includes hands-on model work: supervised fine-tuning and RL on open-source vision-language models, using the proprietary driving data Nuro has collected across years of real-world autonomous operation. You would have a labeling workforce available to you, which means you can specify the data you need rather than making do with what exists. Very few researchers get to run this loop, form a hypothesis about model behavior, commission the exact data to test it, post-train, and evaluate against real driving performance. The fungibility of frontier models is precisely why this matters: the weights are rentable, the data and the labeling capacity behind them are not.

Read the frontier and convert it into experiments. The research frontier moves weekly, and most of it is noise. You are the person who reads it, separates results that will hold from those that will not replicate, and turns the ones that matter into live experiments against our workloads. This is a deliverable, not background reading — our architecture should reflect the best of what is publicly known within weeks of it becoming known.

Test-time scaling. We believe the largest near-term gains come from how inference compute is spent, not from which model is called: reasoning budgets, sampling and search strategies, verifier-guided selection, when to escalate and when to stop. Each buys accuracy at a price, and the exchange rate differs by task. Mapping where additional inference compute pays and where a verifier beats a bigger model is a core agenda here, and it directly determines how we spend a real budget.

Across all of this, we want someone fluent at four levels: frontier models and how they are built; frontier model evaluation and where measurements lie; frontier AI system evaluation, an agent is a model plus tools, memory, retries, verifiers, and a human who accepts or rejects the output, and the compound system fails in ways none of its parts do; and real-world impact, quantified, with the integrity to say when the numbers do not support the story.

About the Work

You will work in close partnership with the engineer who builds the platform this runs on: they make it run safely at scale, you determine what it should be doing and whether it worked.

What You Might Own in Your First Two Quarters

  • Establish the evaluation foundation for the agent fleet we already run: eval data sources, task suites, noise floors, and the statistical standard the team uses to accept or reject a change.
  • Take one high-volume workflow from unmeasured to automatically hill-climbing, end to end, as the template the rest of the system follows.
  • Run a first post-training experiment on an open-weight VLM against our driving data, and establish whether the result justifies the pipeline.
  • Put a defensible number on what the platform is worth: which workflows improved, by how much, with what confidence.

About You

The person we are looking for probably has an engineering background and strong research taste, in that order of surprise. Researchers who can ship are common enough to name; researchers who can ship, judge a literature, and care more about scaled impact than authorship are not. If you have been frustrated by how long it takes good research to reach anything real, this role is the correction.

  • Graduate degree in CS, ML, statistics, or a related field, or equivalent research experience. We care about demonstrated research judgment, not credentials.
  • Fluent in the current literature and able to judge it. You read papers continuously, can tell a real result from a well-marketed one, and have opinions about which recent directions are overrated.
  • Deep understanding of how LLMs work — pretraining through the post-training stack, and what actually happens at inference. You reason from mechanism, not just from published numbers.
  • Firm grasp of the full evaluation pipeline: sourcing eval data, constructing the loop, automating the climb. Having done all three for a real system, rather than one in isolation, is the strongest signal for this role.
  • Rigorous experimentalist. You design experiments that can fail, you understand variance and power, and you are comfortable saying an intervention didn't work.
  • Hands-on post-training experience — SFT and RL, ideally on open-weight models — including the data curation and evaluation work required to know whether it actually helped. Vision-language model experience is a strong plus.
  • A real engineering background. Strong Python, comfortable with production systems and data, able to stand up the infrastructure your own experiment needs.
  • Direct experience with LLM agent systems — building them, evaluating them, or studying why they fail.
  • You measure yourself in impact and in weeks. You want your work in front of hundreds of engineers this quarter.

Bonus Points

  • Published or applied work in agent evaluation, reasoning, test-time compute, RL, or verification.
  • Experience with multimodal or vision-language models, and with data curation at scale.
  • Online experimentation in production: A/B testing, causal inference from observational data, offline-to-online correlation.
  • Experience building evaluation harnesses, task suites, or LLM-as-judge systems, including their failure modes.
  • Familiarity with autonomous systems, safety cases, or verification-gated deployment.

At Nuro, your base pay is one part of your total compensation package. For this position, the reasonably expected base pay range is between $193,930 and $352,290 for the level at which this job has been scoped. Your base pay will depend on several factors, including your experience, qualifications, education, location, and skills. In the event that you are considered for a different level, a higher or lower pay range would apply. This position is also eligible for an annual performance bonus, equity, and a competitive benefits package.

At Nuro, we celebrate differences and are committed to a diverse workplace that fosters inclusion and psychological safety for all employees. Nuro is proud to be an equal opportunity employer and expressly prohibits any form of workplace discrimination based on race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other legally protected characteristics.

Prêt à postuler chez Nuro ?
Postuler chez Nuro

Comment se compare ce salaire pour Researcher

Ce poste paie $273,110/yrau-dessus de la fourchette habituelle pour les postes Researcher.

$80,000 la médiane $170,800 $300,000

Fourchette typique $104,492–$224,838/yr, à partir de 372 annonces Researcher comparables sur JobsRadar (rémunération annualisée en USD). Voir les aperçus de salaire pour Researcher →

Emplois similaires

Protege
Applied Healthcare Researcher
Protege
⚡ Postuler tôt Monde entier
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Illumio
Staff Threat Researcher
Illumio
⚡ Postuler tôt Sunnyvale, California - HQ Sur site $194,000–$233,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
HU
Senior Design Researcher
Hudl
⚡ Postuler tôt Austin, TX, United States; Chi... Sur site $112,000–$187,000
● Nouveau 👁 Vu ✓ Postulé il y a 6 h
DoorDash USA
Associate Manager, AI Research Lab Strategy & Operations
DoorDash USA
⚡ Postuler tôt San Francisco, CA; Sunnyvale,... Sur site $124,000–$155,000
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
DoorDash USA
Member of Technical Staff, Lead Researcher
DoorDash USA
⚡ Postuler tôt San Francisco, CA; Sunnyvale,... Sur site $203,500–$299,300
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Pulse Biosciences
Sr. Clinical Research Associate (Soft Tissue Ablation)
Pulse Biosciences
⚡ Postuler tôt US Remote · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Control Risks
Researcher, Citizenship By Investment and Third-Party Due Diligence
Control Risks
⚡ Postuler tôt Madrid, Community of Madrid, S... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 8 h
SR
Quantitative Researcher Intern
Scientech Research LLC
⚡ Postuler tôt New Jersey Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 9 h
SR
Quantitative Researcher Intern
Scientech Research LLC
⚡ Postuler tôt Shanghai Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 9 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Nuro

Voir tous les emplois chez Nuro →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit