Jobs Companies Perplexity Member of Technical Staff (AI Infrastructure Engineer)

Sobre esta vaga de Member of Technical Staff (AI Infrastructure Engineer) na Perplexity

Perplexity · Presencial · San Francisco

We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads

  • Manage and optimize Slurm-based HPC environments for distributed training of large language models

  • Develop robust APIs and orchestration systems for both training pipelines and inference services

  • Implement resource scheduling and job management systems across heterogeneous compute environments

  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm

  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services

  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Qualifications

  • Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management

  • Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization

  • Experience with deploying and managing distributed training systems at scale

  • Deep understanding of container orchestration and distributed systems architecture

  • High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)

  • Experience managing GPU clusters and optimizing compute resource utilization

Required Skills

  • Expert-level Kubernetes administration and YAML configuration management

  • Proficiency with Slurm job scheduling, resource management, and cluster configuration

  • Python and C++ programming with focus on systems and infrastructure automation

  • Hands-on experience with ML frameworks such as PyTorch in distributed training contexts

  • Strong understanding of networking, storage, and compute resource management for ML workloads

  • Experience developing APIs and managing distributed systems for both batch and real-time workloads

  • Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred Skills

  • Experience with Kubernetes operators and custom controllers for ML workloads

  • Advanced Slurm administration including multi-cluster federation and advanced scheduling policies

  • Familiarity with GPU cluster management and CUDA optimization

  • Experience with other ML frameworks like TensorFlow or distributed training libraries

  • Background in HPC environments, parallel computing, and high-performance networking

  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices

  • Experience with container registries, image optimization, and multi-stage builds for ML workloads

Required Experience

  • Demonstrated experience managing large-scale Kubernetes deployments in production environments

  • Proven track record with Slurm cluster administration and HPC workload management

  • Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure

  • Experience supporting both long-running training jobs and high-availability inference services

  • Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Pronto para se candidatar à Perplexity?
Candidatar-se à Perplexity

Como este salário de Platform Engineer se compara

Esta vaga paga $312,500/yracima da faixa típica para vagas de Platform Engineer.

$177,500 a mediana $234,250 $360,000

Faixa típica $200,551–$285,500/yr, com base em 457 vagas de Platform Engineer comparáveis na JobsRadar (pagamento anualizado em USD). Ver insights salariais de Platform Engineer →

Vagas semelhantes

a16z
Partner 20, Staff Security Engineer, AI & Security Platform
a16z
⚡ Candidate-se cedo San Francisco, California, Uni... Presencial $243,000–$284,000
● Nova 👁 Vista ✓ Candidatada há 5h
LangChain
Senior Fullstack Engineer, AI Observability & Evals Platform
LangChain
⚡ Candidate-se cedo San Francisco, CA Presencial $175,000–$240,000
● Nova 👁 Vista ✓ Candidatada há 10h
SA
Principal Software Engineer, Platform Services (PMTS)
Salesforce
⚡ Candidate-se cedo California - San Francisco Presencial $237,700–$344,700
● Nova 👁 Vista ✓ Candidatada há 14h
MI
Software Engineer - Member of Technical Staff (Platform)
Mithril
⚡ Candidate-se cedo Palo Alto / San Francisco Bay... Presencial $170,000–$230,000
● Nova 👁 Vista ✓ Candidatada há 16h
Vapi
Member of Technical Staff, Infrastructure Engineer
Vapi
⚡ Candidate-se cedo San Francisco Híbrido $280,000–$314,000
● Nova 👁 Vista ✓ Candidatada há 16h
Redwood Materials
Infrastructure Software Engineer, Energy Storage
Redwood Materials
⚡ Candidate-se cedo San Francisco, California, Uni... Presencial $180,000–$237,500
● Nova 👁 Vista ✓ Candidatada há 17h
Wispr Flow
Platform Engineer, Billing Systems
Wispr Flow
⚡ Candidate-se cedo San Francisco Presencial $220,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 18h
Wispr Flow
Platform Engineer, Infrastructure
Wispr Flow
⚡ Candidate-se cedo San Francisco Presencial $220,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 18h
Wispr Flow
Platform Engineer, Product Security
Wispr Flow
⚡ Candidate-se cedo San Francisco Presencial $220,000–$350,000
● Nova 👁 Vista ✓ Candidatada há 18h

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na Perplexity

Ver todas as vagas na Perplexity →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis