Jobs › Companies › NVIDIA › Senior DevOps Engineer, AIOps

À propos de ce poste Senior DevOps Engineer, AIOps chez NVIDIA

NVIDIA · Sur site · Israel, Raanana

NVIDIA is powering the world’s most advanced AI factories, where resilient infrastructure is essential to keep accelerated computing environments running at scale. The Agentic AIOps team is building a mission-critical observability and prediction platform - delivered as both a high-scale SaaS solution and a robust on-premises deployment for NVIDIA’s largest enterprise customers. 


As a Senior DevOps Engineer, you’ll help turn agentic AI capabilities for diagnosing and troubleshooting network and GPU infrastructure into secure, scalable, production-ready services. This role stands out through its end-to-end ownership across cloud and customer-managed environments, close partnership with software and AI engineers, and direct influence on the reliability of NVIDIA’s AI infrastructure. 


What You'll Be Doing: 

  • Own the DevOps, infrastructure, security, release, and reliability lifecycle - from development environments and CI/CD through deployment, production readiness, and sustained operations. 
  • Build and operate Kubernetes environments and Helm-based deployments for a Python, FastAPI, Node.js, and React microservices platform across SaaS and on-premises footprints. 
  • Engineer GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory. 
  • Automate infrastructure provisioning, configuration, upgrades, and routine operational workflows to accelerate delivery and improve engineering productivity. 
  • Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage with disciplined capacity planning, backups, recovery testing, and safe migrations. 
  • Strengthen release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery plans for active workflows. 
  • Deliver actionable observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening. 
  • Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance, resource efficiency, and customer outcomes. 

 

What We Need to See: 

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience. 
  • 5+ years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices. 
  • Strong hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting. 
  • Strong Linux administration skills and proficiency in Python and Bash for automation, plus experience with infrastructure as code and configuration tooling such as Terraform and Ansible. 
  • Experience building and maintaining CI/CD pipelines, including runners, container registries, artifact management, automated quality gates, and secure release practices. 
  • Practical experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting. 
  • Strong networking and observability fundamentals across TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and actionable alerting. 
  • Sound understanding of secure infrastructure operations and incident response, with demonstrated ownership, cross-functional collaboration, and prioritization in an evolving environment. 

 

Ways To Stand Out From the Crowd: 

  • Experience operating AI applications, agent platforms, or LLM services, including monitoring latency, failures, token usage, and cost. 
  • Familiarity with Temporal, LangGraph, Model Context Protocol (MCP), Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage. 
  • Deep experience with OpenTelemetry instrumentation and collectors, Datadog APM, or Prometheus/Grafana. 
  • Experience with self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, or GPU clusters and AI data centers. 
  • Experience building reproducible AMD64 and ARM64 container images, optimizing BuildKit pipelines, and securing the software supply chain. 

 

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are passionate about building mission-critical systems at the frontier of AI infrastructure, we want to hear from you. 

 

#LI-Hybrid

Prêt à postuler chez NVIDIA ?
Postuler chez NVIDIA

À propos de NVIDIA

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA .

Voir tous les emplois chez NVIDIA →

Emplois similaires

NVIDIA
DevOps Engineer
NVIDIA
⚡ Postuler tôt Israel, Raanana Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
NVIDIA
Senior System Software Engineer - Signing Platform
NVIDIA
⚡ Postuler tôt Israel, Raanana Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 sem.
NVIDIA
Senior DevOps Engineer
NVIDIA
⚡ Postuler tôt Israel, Raanana Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 mois
Octal Philippines Inc.
Sr. DevOps Engineer
Octal Philippines Inc.
⚡ Postuler tôt Taguig City, National Capital... Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 45 min
Quast Ltd
Reliability & Maintainability Engineer - Mobile Fires Platform
Quast Ltd
⚡ Postuler tôt Stoke Gifford, England, United... Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
TrustEngine
Lead Database & Data Platform Engineer
TrustEngine
⚡ Postuler tôt United States · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Turquoise
Platform Operations Engineer
Turquoise
⚡ Postuler tôt Remote · lieu restreint $153,000–$170,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Suno
Senior/Staff Software Engineer - Commerce Platform
Suno
⚡ Postuler tôt NYC Sur site $228,960–$363,510
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Scribe
Senior DevOps Engineer
Scribe
⚡ Postuler tôt Remote (PST) Hybride $150,000–$240,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez NVIDIA

Voir tous les emplois chez NVIDIA →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit