Jobs Companies NVIDIA Senior Site Reliability Engineer in Test, SDET

Sobre este puesto de Senior Site Reliability Engineer in Test, SDET en NVIDIA

NVIDIA · Presencial · China, Shanghai

At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing. Our team in Shanghai, China is looking for a Senior Site Reliability Engineer focused on Test Environment Management to join us. This is an opportunity to build and operate highly reliable test infrastructure, CI/CD systems, and environments that power validation of NVIDIA enterprise offerings. If you are passionate about reliability engineering, test infrastructure, and AI-scale systems, this role is for you.


What you’ll be doing

  • Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure that support validation of NVIDIA enterprise offerings.

  • Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps) — including pipeline design, reliability, performance, and progressive delivery of test workloads.

  • Manage Software Bills of Materials (SBOMs): generation, continuous monitoring, vulnerability correlation, policy enforcement, and integration into CI/CD and release gate..

  • Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments (Kubernetes-based and hybrid) with strong emphasis on isolation, reproducibility, and rapid recovery.

  • Define and drive reliability practices for test systems: SLIs/SLOs/error budgets for test environments and pipelines, toil reduction, chaos/resilience testing of infrastructure, and automated remediation.

  • Collaborate closely with development and platform teams to triage environment and pipeline failures, perform root-cause analysis, verify fixes, and continuously harden test infrastructure.

  • Apply AI/ML/Agentic techniques and internal tools to accelerate environment provisioning, flaky-test detection, capacity planning, anomaly detection, and overall Quality Assurance velocity.

What we need to see

  • MS or PhD in Computer Science, related field, or equivalent experience with 8+ years of professional experience in Site Reliability Engineering, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.

  • Strong proficiency with Linux, shell scripting, and Python (or equivalent automation languages).

  • Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions, practical experience with ArgoCD (or equivalent GitOps tooling) for CD of applications and infrastructure. Solid background in containerization and orchestration (Docker, Kubernetes) and virtualization technologies.

  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination, and reliability engineering for complex distributed systems.

  • Experience building and operating test environments (ephemeral, multi-tenant, or production-like) with focus on reliability, isolation, and rapid turnaround.

  • Strong knowledge of QA principles and how test infrastructure enables high-quality software delivery.

  • Comfort working with AI/LLM-related workloads and toolings. Excellent problem-solving skills, clear written and verbal communication, and the ability to collaborate across engineering teams.

  • Self-motivated, proactive, and passionate about learning new technology at scale.

Ways to stand out from the crowd

  • Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.

  • Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno, etc.).

  • Hands-on work with NVIDIA GPU hardware, multi-GPU environments, or accelerated computing infrastructure. Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses.

  • Track record of applying AI/observability techniques to detect flaky tests, optimize environment utilization, or automate root-cause analysis.

  • Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.

NVIDIA is widely considered one of the technology world’s most desirable employers. We have some of the most brilliant people on the planet working for us. If you’re creative, autonomous, and excited about making test infrastructure as reliable as the products it validates, we want to hear from you!

¿Listo para postularte en NVIDIA?
Postúlate en NVIDIA

Sobre NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Ver todos los empleos en NVIDIA →

Empleos similares

Optiver
Production Operations Engineer (Trading Systems / SRE / Application Support)
Optiver
⚡ Postúlate pronto Shanghai, China Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1 mes
Dyson
Senior Reliability Engineer
Dyson
⚡ Postúlate pronto China - Shanghai Office Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1 mes
Air Liquide
Electrical / Reliability Engineer (intermediate) - EL
Air Liquide
⚡ Postúlate pronto China, Shanghai Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1 mes
HI
Principal Engineer, Product Reliability
HARMAN International
⚡ Postúlate pronto Suzhou - Jiangsu, China Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1 mes
Roku
Senior Software Engineer, MLOps/SRE
Roku
⚡ Postúlate pronto Bengaluru, India Presencial
● Nuevo 👁 Visto ✓ Postulado hace 6h
Roku
Senior Software Engineer,  SRE
Roku
⚡ Postúlate pronto Bengaluru, India Presencial
● Nuevo 👁 Visto ✓ Postulado hace 6h
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Postúlate pronto San Mateo, CA, United States Presencial $243,290–$295,250
● Nuevo 👁 Visto ✓ Postulado hace 11h
Genesys
Senior Operations Reliability Engineer - IAM
Genesys
⚡ Postúlate pronto Ontario, Canada Presencial
● Nuevo 👁 Visto ✓ Postulado hace 12h
LSEG
Senior Manager, Site Reliability Engineering
LSEG
⚡ Postúlate pronto IND-BLR-Divyasree Technopolis Presencial
● Nuevo 👁 Visto ✓ Postulado hace 13h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en NVIDIA

Ver todos los empleos en NVIDIA →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis