Jobs Companies Perplexity Member of Technical Staff (AI Infrastructure Engineer)

About this Member of Technical Staff (AI Infrastructure Engineer) role at Perplexity

Perplexity · Onsite · San Francisco

We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams to build, deploy, and optimize our large-scale AI training and inference clusters

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads

  • Manage and optimize Slurm-based HPC environments for distributed training of large language models

  • Develop robust APIs and orchestration systems for both training pipelines and inference services

  • Implement resource scheduling and job management systems across heterogeneous compute environments

  • Benchmark system performance, diagnose bottlenecks, and implement improvements across both training and inference infrastructure

  • Build monitoring, alerting, and observability solutions tailored to ML workloads running on Kubernetes and Slurm

  • Respond swiftly to system outages and collaborate across teams to maintain high uptime for critical training runs and inference services

  • Optimize cluster utilization and implement autoscaling strategies for dynamic workload demands

Qualifications

  • Strong expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management

  • Hands-on experience with Slurm workload management, including job scheduling, resource allocation, and cluster optimization

  • Experience with deploying and managing distributed training systems at scale

  • Deep understanding of container orchestration and distributed systems architecture

  • High level familiarity with LLM architecture and training processes (Multi-Head Attention, Multi/Grouped-Query, distributed training strategies)

  • Experience managing GPU clusters and optimizing compute resource utilization

Required Skills

  • Expert-level Kubernetes administration and YAML configuration management

  • Proficiency with Slurm job scheduling, resource management, and cluster configuration

  • Python and C++ programming with focus on systems and infrastructure automation

  • Hands-on experience with ML frameworks such as PyTorch in distributed training contexts

  • Strong understanding of networking, storage, and compute resource management for ML workloads

  • Experience developing APIs and managing distributed systems for both batch and real-time workloads

  • Solid debugging and monitoring skills with expertise in observability tools for containerized environments

Preferred Skills

  • Experience with Kubernetes operators and custom controllers for ML workloads

  • Advanced Slurm administration including multi-cluster federation and advanced scheduling policies

  • Familiarity with GPU cluster management and CUDA optimization

  • Experience with other ML frameworks like TensorFlow or distributed training libraries

  • Background in HPC environments, parallel computing, and high-performance networking

  • Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices

  • Experience with container registries, image optimization, and multi-stage builds for ML workloads

Required Experience

  • Demonstrated experience managing large-scale Kubernetes deployments in production environments

  • Proven track record with Slurm cluster administration and HPC workload management

  • Previous roles in SRE, DevOps, or Platform Engineering with focus on ML infrastructure

  • Experience supporting both long-running training jobs and high-availability inference services

  • Ideally, 3-5 years of relevant experience in ML systems deployment with specific focus on cluster orchestration and resource management

Ready to apply to Perplexity?
Apply to Perplexity

How this Platform Engineer salary compares

This role pays $312,500/yrabove the typical range for Platform Engineer roles.

$177,500 median $234,500 $360,000

Typical range $200,704–$285,500/yr, from 456 comparable Platform Engineer listings on JobsRadar (pay annualized to USD). See Platform Engineer salary insights →

Similar jobs

Redwood Materials
Infrastructure Software Engineer, Energy Storage
Redwood Materials
⚡ Apply early San Francisco, California, Uni... Onsite $180,000–$237,500
● New 👁 Seen ✓ Applied 6h ago
Doctronic
Senior Platform Engineer
Doctronic
⚡ Apply early San Francisco Onsite $180,000–$240,000
● New 👁 Seen ✓ Applied 13h ago
Firecrawl
Backend Infrastructure Engineer
Firecrawl
⚡ Apply early San Francisco HQ; Toronto Hub Onsite $235,000–$260,000
● New 👁 Seen ✓ Applied 21h ago
Simple AI
Infrastructure Engineer
Simple AI
⚡ Apply early San Francisco Onsite
● New 👁 Seen ✓ Applied 1d ago
a16z
Partner 20, Staff Security Engineer, AI & Security Platform
a16z
⚡ Apply early San Francisco, California, Uni... Hybrid $243,000–$284,000
● New 👁 Seen ✓ Applied 1d ago
LangChain
Senior Fullstack Engineer, AI Observability & Evals Platform
LangChain
⚡ Apply early San Francisco, CA Onsite $175,000–$240,000
● New 👁 Seen ✓ Applied 2d ago
SA
Principal Software Engineer, Platform Services (PMTS)
Salesforce
⚡ Apply early California - San Francisco Onsite $237,700–$344,700
● New 👁 Seen ✓ Applied 2d ago
MI
Software Engineer - Member of Technical Staff (Platform)
Mithril
⚡ Apply early Palo Alto / San Francisco Bay... Hybrid $170,000–$230,000
● New 👁 Seen ✓ Applied 2d ago
Vapi
Member of Technical Staff, Infrastructure Engineer
Vapi
⚡ Apply early San Francisco Hybrid $280,000–$314,000
● New 👁 Seen ✓ Applied 2d ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Perplexity

See all jobs at Perplexity →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free