Jobs Companies NVIDIA Senior Staff SRE – Compute Platform

À propos de ce poste Senior Staff SRE – Compute Platform chez NVIDIA

NVIDIA · Sur site · India, Bengaluru

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations.

Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.

What you’ll be doing:

  • Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.

  • Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.

  • Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.

  • Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.

  • Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.

What we need to see:

  • BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.

  • Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.

  • Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.

  • Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.

  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.

  • Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.

  • Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.

Ways to stand out from the crowd:

  • Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.

  • Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.

  • Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.

  • Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.

  • Demonstrated delivery of complex, high-impact infrastructure projects.

Prêt à postuler chez NVIDIA ?
Postuler chez NVIDIA

À propos de NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Voir tous les emplois chez NVIDIA →

Emplois similaires

Roku
Senior Software Engineer, MLOps/SRE
Roku
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Roku
Senior Software Engineer,  SRE
Roku
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Fivetran
Staff Site Reliability Engineer
Fivetran
⚡ Postuler tôt Bengaluru, Karnataka, India, A... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
NVIDIA
Senior Staff Site Reliability Engineer
NVIDIA
⚡ Postuler tôt India, Bengaluru Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Aviatrix
Senior MTS - SRE
Aviatrix
⚡ Postuler tôt Bengaluru, Karnataka, India (A... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Okta
Manager- Site Reliability Engineer
Okta
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 j
OneTrust
Senior Staff Software Engineer-SRE
OneTrust
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 j
Harness
Staff Software Engineer – AI SRE
Harness
⚡ Postuler tôt Bengaluru, Karnataka, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
NVIDIA
Senior Site Reliability Engineering, Storage
NVIDIA
⚡ Postuler tôt India, Bengaluru Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez NVIDIA

Voir tous les emplois chez NVIDIA →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit