Jobs Companies Perplexity Member of Technical Staff (AI Infrastructure Engineer)

À propos de ce poste Member of Technical Staff (AI Infrastructure Engineer) chez Perplexity

Perplexity · Sur site · San Francisco

We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads

  • Manage and optimize Slurm-based HPC environments for distributed training of large language models

  • Develop robust APIs and orchestration systems for both training pipelines and inference services

  • Implement resource scheduling and job management systems across heterogeneous compute environments

  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm

  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services

  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Qualifications

  • Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management

  • Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization

  • Experience with deploying and managing distributed training systems at scale

  • Deep understanding of container orchestration and distributed systems architecture

  • High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)

  • Experience managing GPU clusters and optimizing compute resource utilization

Required Skills

  • Expert-level Kubernetes administration and YAML configuration management

  • Proficiency with Slurm job scheduling, resource management, and cluster configuration

  • Python and C++ programming with focus on systems and infrastructure automation

  • Hands-on experience with ML frameworks such as PyTorch in distributed training contexts

  • Strong understanding of networking, storage, and compute resource management for ML workloads

  • Experience developing APIs and managing distributed systems for both batch and real-time workloads

  • Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred Skills

  • Experience with Kubernetes operators and custom controllers for ML workloads

  • Advanced Slurm administration including multi-cluster federation and advanced scheduling policies

  • Familiarity with GPU cluster management and CUDA optimization

  • Experience with other ML frameworks like TensorFlow or distributed training libraries

  • Background in HPC environments, parallel computing, and high-performance networking

  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices

  • Experience with container registries, image optimization, and multi-stage builds for ML workloads

Required Experience

  • Demonstrated experience managing large-scale Kubernetes deployments in production environments

  • Proven track record with Slurm cluster administration and HPC workload management

  • Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure

  • Experience supporting both long-running training jobs and high-availability inference services

  • Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Prêt à postuler chez Perplexity ?
Postuler chez Perplexity

Comment se compare ce salaire pour Platform Engineer

Ce poste paie $312,500/yrau-dessus de la fourchette habituelle pour les postes Platform Engineer.

$178,500 la médiane $226,000 $345,150

Fourchette typique $200,000–$275,500/yr, à partir de 414 annonces Platform Engineer comparables sur JobsRadar (rémunération annualisée en USD). Voir les aperçus de salaire pour Platform Engineer →

Emplois similaires

OpenAI
Platform Engineer, Forward Deployed Engineering (FDE) -SF
OpenAI
⚡ Postuler tôt San Francisco Hybride $230,000–$385,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Handshake
Platform Engineer, Handshake AI Enterprise
Handshake
⚡ Postuler tôt San Francisco, CA Sur site $155,000–$300,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 h
Stitch Fix
ML Platform Engineer
Stitch Fix
⚡ Postuler tôt Remote, USA · lieu restreint $136,000–$167,000
● Nouveau 👁 Vu ✓ Postulé il y a 7 h
Anthropic
Staff Software Engineer, Environments Infrastructure
Anthropic
⚡ Postuler tôt San Francisco, CA | New York C... Sur site $405,000–$605,000
● Nouveau 👁 Vu ✓ Postulé il y a 9 h
Baton (A Ryder Technology Lab)
Staff Software Engineer - Infrastructure, Data Platform
Baton (A Ryder Technology Lab)
⚡ Postuler tôt San Francisco, California, Uni... Sur site $260,000–$350,000
● Nouveau 👁 Vu ✓ Postulé il y a 10 h
Databricks
Staff Software Engineer - AI Platform (NYC)
Databricks
⚡ Postuler tôt New York City, New York Sur site $190,900–$253,750
● Nouveau 👁 Vu ✓ Postulé il y a 10 h
Databricks
Sr Software Engineer, Infrastructure
Databricks
⚡ Postuler tôt San Francisco, California Sur site $159,900–$219,900
● Nouveau 👁 Vu ✓ Postulé il y a 10 h
Databricks
Senior Software Engineer, Compute Infrastructure
Databricks
⚡ Postuler tôt Mountain View, California; San... Sur site $164,200–$205,200
● Nouveau 👁 Vu ✓ Postulé il y a 10 h
Salient
Infrastructure Engineer
Salient
⚡ Postuler tôt SF Headquarters Sur site $200,000–$300,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Perplexity

Voir tous les emplois chez Perplexity →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit