Jobs Companies Perplexity Member of Technical Staff (AI Infrastructure Engineer)

Sobre esta vaga de Member of Technical Staff (AI Infrastructure Engineer) na Perplexity

Perplexity · Presencial · San Francisco

We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads

  • Manage and optimize Slurm-based HPC environments for distributed training of large language models

  • Develop robust APIs and orchestration systems for both training pipelines and inference services

  • Implement resource scheduling and job management systems across heterogeneous compute environments

  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm

  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services

  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Qualifications

  • Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management

  • Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization

  • Experience with deploying and managing distributed training systems at scale

  • Deep understanding of container orchestration and distributed systems architecture

  • High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)

  • Experience managing GPU clusters and optimizing compute resource utilization

Required Skills

  • Expert-level Kubernetes administration and YAML configuration management

  • Proficiency with Slurm job scheduling, resource management, and cluster configuration

  • Python and C++ programming with focus on systems and infrastructure automation

  • Hands-on experience with ML frameworks such as PyTorch in distributed training contexts

  • Strong understanding of networking, storage, and compute resource management for ML workloads

  • Experience developing APIs and managing distributed systems for both batch and real-time workloads

  • Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred Skills

  • Experience with Kubernetes operators and custom controllers for ML workloads

  • Advanced Slurm administration including multi-cluster federation and advanced scheduling policies

  • Familiarity with GPU cluster management and CUDA optimization

  • Experience with other ML frameworks like TensorFlow or distributed training libraries

  • Background in HPC environments, parallel computing, and high-performance networking

  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices

  • Experience with container registries, image optimization, and multi-stage builds for ML workloads

Required Experience

  • Demonstrated experience managing large-scale Kubernetes deployments in production environments

  • Proven track record with Slurm cluster administration and HPC workload management

  • Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure

  • Experience supporting both long-running training jobs and high-availability inference services

  • Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Pronto para se candidatar à Perplexity?
Candidatar-se à Perplexity

Como este salário de Platform Engineer se compara

Esta vaga paga $312,500/yracima da faixa típica para vagas de Platform Engineer.

$178,500 a mediana $226,000 $345,150

Faixa típica $200,000–$275,500/yr, com base em 414 vagas de Platform Engineer comparáveis na JobsRadar (pagamento anualizado em USD). Ver insights salariais de Platform Engineer →

Vagas semelhantes

OpenAI
Platform Engineer, Forward Deployed Engineering (FDE) -SF
OpenAI
⚡ Candidate-se cedo San Francisco Híbrido $230,000–$385,000
● Nova 👁 Vista ✓ Candidatada há 4h
Handshake
Platform Engineer, Handshake AI Enterprise
Handshake
⚡ Candidate-se cedo San Francisco, CA Presencial $155,000–$300,000
● Nova 👁 Vista ✓ Candidatada há 4h
Stitch Fix
ML Platform Engineer
Stitch Fix
⚡ Candidate-se cedo Remote, USA · local restrito $136,000–$167,000
● Nova 👁 Vista ✓ Candidatada há 6h
Anthropic
Staff Software Engineer, Environments Infrastructure
Anthropic
⚡ Candidate-se cedo San Francisco, CA | New York C... Presencial $405,000–$605,000
● Nova 👁 Vista ✓ Candidatada há 8h
Baton (A Ryder Technology Lab)
Staff Software Engineer - Infrastructure, Data Platform
Baton (A Ryder Technology Lab)
⚡ Candidate-se cedo San Francisco, California, Uni... Presencial $260,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 9h
Databricks
Staff Software Engineer - AI Platform (NYC)
Databricks
⚡ Candidate-se cedo New York City, New York Presencial $190,900–$253,750
● Nova 👁 Vista ✓ Candidatada há 9h
Databricks
Sr Software Engineer, Infrastructure
Databricks
⚡ Candidate-se cedo San Francisco, California Presencial $159,900–$219,900
● Nova 👁 Vista ✓ Candidatada há 9h
Databricks
Senior Software Engineer, Compute Infrastructure
Databricks
⚡ Candidate-se cedo Mountain View, California; San... Presencial $164,200–$205,200
● Nova 👁 Vista ✓ Candidatada há 9h
Salient
Infrastructure Engineer
Salient
⚡ Candidate-se cedo SF Headquarters Presencial $200,000–$300,000
● Nova 👁 Vista ✓ Candidatada há 4h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na Perplexity

Ver todas as vagas na Perplexity →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis