Jobs Companies NVIDIA Senior Staff Site Reliability Engineer

À propos de ce poste Senior Staff Site Reliability Engineer chez NVIDIA

NVIDIA · Sur site · India, Bengaluru

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

What you'll be doing:

  • Lead the technical strategy and roadmap for large-scale, multi-functional SRE initiatives that boost reliability, scalability, and developer efficiency throughout enterprise systems.

  • Design and build resilient distributed systems that power NVIDIA's next-generation AI-powered enterprise products and services, including transforming legacy applications and database systems into modern and scalable architectures.

  • Drive automation and observability improvements, using metrics and analytics — including AI workload quality signals and model performance telemetry — to improve performance, reliability, and efficiency.

  • Build LLM-aware monitoring and autonomous incident response pipelines to reduce toil, accelerate MTTR, and evolve on-call operations toward AI-assisted remediation.

  • Work together with Cloud, Platform, Security, and AI/ML groups to develop modern SRE elements and AI-native platform features that guarantee high availability and secure operations.

  • Analyze and run complex systems — including Kubernetes-scale and AI/ML infrastructure challenges — championing standards in system design and incident management.

  • Drive AI-assisted and AI-first engineering practices across the organization, mentoring engineers to adopt agentic development workflows, coding agents, and LLM-powered tooling in their day-to-day work.

What we need to see:

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.

  • BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience.

  • Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code and distributed systems debugging.

  • Experience with infrastructure-as-code tooling such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.

  • Solid understanding of OpenTelemetry or other observability implementations at scale, including observability build for AI workloads (model performance, drift detection, and AI quality signals).

  • Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP), including bare-metal and GPU-accelerated infrastructure.

  • Outstanding problem-solving, communication, and collaboration skills, with the ability to influence across technical and interpersonal boundaries.

Ways to stand out from the crowd:

  • Experience with Public Cloud or large-scale automation systems.

  • Capability to steer technical strategy and achieve quantifiable reliability results in complex, multi-team settings.

  • Experience building or operating agentic AI platforms — autonomous or semi-autonomous — including LLM toolchains, agent orchestration frameworks (e.g., LangGraph, AutoGen), or automated runbook and toil-reduction processes.

  • A strong sense of ownership, curiosity, and innovation — and turn challenges into opportunities.

NVIDIA leads the charge in innovative breakthroughs in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, functions as the visual cortex of today’s computers and forms the core of our products and services. Our work opens new realms to explore, encourages outstanding creativity and discovery, and powers inventions once thought of as science fiction — from artificial intelligence to autonomous systems. NVIDIA is searching for outstanding talent like you to help us advance the next wave of artificial intelligence!

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com

Prêt à postuler chez NVIDIA ?
Postuler chez NVIDIA

À propos de NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Voir tous les emplois chez NVIDIA →

Emplois similaires

Roku
Senior Software Engineer, MLOps/SRE
Roku
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 h
Roku
Senior Software Engineer,  SRE
Roku
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 h
Fivetran
Staff Site Reliability Engineer
Fivetran
⚡ Postuler tôt Bengaluru, Karnataka, India, A... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Wells Fargo
Lead Infrastructure Engineer (Site Reliability Engineering) – Workplace Technology
Wells Fargo
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Fiserv
Tech Lead, Infrastructure Engineering(Cloud SRE L2)
Fiserv
⚡ Postuler tôt Bengaluru - GS, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
NVIDIA
Senior Staff Site Reliability Engineer
NVIDIA
⚡ Postuler tôt India, Bengaluru Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Aviatrix
Senior MTS - SRE
Aviatrix
⚡ Postuler tôt Bengaluru, Karnataka, India (A... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Stryker
Senior DevSecOps/Site Reliability Engineer (AWS)
Stryker
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 j
Okta
Manager- Site Reliability Engineer
Okta
⚡ Postuler tôt Bengaluru, India Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 j

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez NVIDIA

Voir tous les emplois chez NVIDIA →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit