Jobs › Companies › Avride › Lead Evaluation Engineer – Scenario Coverage & Datasets (Autonomous Driving)

Sobre esta vaga de Lead Evaluation Engineer – Scenario Coverage & Datasets (Autonomous Driving) na Avride

Avride · Presencial · Austin, Texas

Lead Evaluation Engineer Scenario Coverage & Datasets (Autonomous Driving)

About the team

Avride builds autonomous driving technology for vehicles and delivery robots. Our QA organization measures how well that technology drives, largely in simulation — and an evaluation is only as good as the data behind it. So we build our own datasets, keep track of what they cover, and keep widening that coverage as the technology takes on more of the road.

About the role

You will own test coverage for the autonomous driving stack, end to end. Scenario areas are owned by individual QA engineers. You own the whole: what our evaluation dataset must contain for a release decision to be trustworthy, where the gaps are that nobody working inside a single area can see, and how those gaps keep surfacing without you being the one who finds them.

Coverage is a data problem before it is a testing problem. A scenario that exists on paper is not covered until there are real scenes behind it, and closing that distance is the craft of this role: metric-driven search across driving data, VLM-assisted retrieval and clustering, criticality-based selection, variation in simulation, and structured capture in the field. You will use that toolkit yourself before you ask anyone else to, and turn what works into something the whole QA team can run without you.

You will lead through process first — building it, and delivering through the QA engineers and operational resources already around you — and grow a team of your own as the scope demands it.

The mindset we're hiring for

Most testing work starts from a defined expected output, where the job is to confirm the system produced it. This role is different. The possible driving situations are effectively infinite, and for most of them there is no single right answer to check against.

We're looking for someone who thinks in terms of a space rather than a list: which parts of it we've sampled, which we haven't, what's rare but matters, and how you'd even notice that something is missing. If you've had to decide what to test when you couldn't test everything, and defend that decision, this role will feel familiar.

You don't need autonomous vehicle experience

The AV domain is learnable in a couple of quarters. The coverage instinct is what we can't teach. If you've done this kind of work under a different name, we want to hear from you:

  • Evaluation, validation, or simulation in AV, robotics, or perception
  • Detection engineering or threat detection, where the attack space is unbounded, there's no ground truth per event, and you're always trading false positives against false negatives
  • Search relevance or recommendations, building judged sets with coverage across query types
  • LLM or ML model evaluation, building eval sets and benchmarks and finding what they fail to catch
  • Speech recognition, medical AI, or fraud and risk modeling, curating test sets for the long tail, edge cases, and drift

What you'll do

  • Own the evaluation dataset. What is in it, what is missing, and what it lets us claim about autonomous driving quality.
  • Own the coverage model. Our scenario taxonomy and ODD parameter space have to stay accurate as the technology matures and the operating environment changes. You own that: spotting where the model no longer describes what our vehicles actually meet, and working with the QA engineer who owns each category to keep it current.
  • Close the distance between scenarios and scenes. For each thin area, choose the method that will actually produce the scenes you need — mining, simulation or the field — and know what it costs before you spend it.
  • Set the mining agenda for the QA team. Decide what each area owner should be looking for next, review what comes back, and keep a regular rhythm for collecting their feedback.
  • Be the internal customer for our mining tooling. Use it yourself, turn what you learn into requirements for the development and analytics teams, and keep the feedback loop between QA and engineering running so the tooling keeps pace with what we ask of it.
  • Bring in what works elsewhere. Track how other AV programs, research groups and vendors solve coverage, edge-case curation and data selection, and turn what is worth having into concrete proposals.
  • Work through the teams you depend on. Coverage is only visible once scenes are labeled and only measurable when the right metrics exist, which makes labeling and analytics standing partners rather than occasional ones. Give them clear priorities and specific, actionable feedback, and stay close enough to see early when something you depend on is at risk.
  • Report coverage and readiness. Put them in a form a release decision can be made on — clear about confidence and about blind spots.

What you’ll need

  • 8+ years in software testing, test engineering, validation, or evaluation, including 2+ years owning the test or evaluation strategy for a full system rather than a feature area.
  • A track record of taking on more. You've owned an area end to end, led or mentored engineers, and improved how your team works rather than only executing within it.
  • Coverage as a first-class problem. You can talk about equivalence classes, parameter spaces, risk-based prioritization, and what would have to be true for a dataset to be enough.
  • Hands-on with data. SQL and Python at the level where you pull, join and sanity-check a metrics dataset yourself rather than filing a ticket.
  • Fluent use of LLMs as a working tool — log search, triage, scenario generation, report drafting, quick analysis scripts. 
  • Strong written communication. Your reports are read by engineering leads and by our safety organization.
  • Ability to reason precisely about US road rules and real-world driver behavior — or to get there quickly.

Nice to have

  • Experience hiring QA engineers — screening, interview loops, and hire/no-hire calls you stand behind, whether into your own team or a partner team.
  • Autonomous vehicles, robotics, ADAS — or another safety-critical domain (aerospace, medical devices, rail, industrial automation). Helpful, not required.
  • Familiarity with scenario-based testing frameworks: ODD taxonomy (ISO 34503), scenario-based safety evaluation (ISO 34501/34502), and the SOTIF (ISO 21448) known/unknown × safe/unsafe framing. Not required; you'll pick these up on the job.
  • Experience with data mining, active learning or data-selection loops over large sensor, log, or model-output datasets.
  • Experience specifying internal tooling and working with a platform team.
  • Experience with vehicle dynamics or sensor modalities (LiDAR, radar, cameras).

 

#LS-MS1

Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.

Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email [email protected].

Pronto para se candidatar à Avride?
Candidatar-se à Avride

Vagas semelhantes

Avride
Robot Service Technician – Full Time/Contract
Avride
⚡ Candidate-se cedo University, Mississippi Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Fleet Maintenance & Detailing Associate
Avride
⚡ Candidate-se cedo Dallas, Texas Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Robot Service Technician – Full Time/Contract
Avride
⚡ Candidate-se cedo Arlington, Virginia Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Lead AI Infrastructure Engineer
Avride
⚡ Candidate-se cedo Austin, Texas Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Robot Service Technician – Full Time/Contract (Mon-Fri, PM Shift)
Avride
⚡ Candidate-se cedo Lexington, Kentucky Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Robot Service Technician – Full Time/Contract
Avride
⚡ Candidate-se cedo Philadelphia, Pennsylvania Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Facilities Technician
Avride
⚡ Candidate-se cedo Austin, Texas Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Fleet Maintenance & Detailing Associate
Avride
⚡ Candidate-se cedo Austin, Texas Presencial
● Nova 👁 Vista ✓ Candidatada há 3h
Avride
Lead Test Engineer AV Safety Assurance
Avride
⚡ Candidate-se cedo Austin, Texas Presencial
● Nova 👁 Vista ✓ Candidatada há 3h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na Avride

Ver todas as vagas na Avride →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis