Jobs Companies Inferact Member of Technical Staff, Site Reliability Engineer

Sobre este puesto de Member of Technical Staff, Site Reliability Engineer en Inferact

Inferact · Presencial · San Francisco

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.

You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.

  • Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.

  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.

  • Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.

  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.

  • Ability to design operationally simple systems and identify likely failure modes before launch.

  • Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.

Preferred qualifications:

  • Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.

  • Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.

  • Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.

  • Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.

  • Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.

Bonus points if you have:

  • Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.

  • Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.

  • Built automation that reduced toil, improved recovery time, or prevented repeat incidents.

  • Led incident response for severe outages with clear communication across engineering and leadership.

  • Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: We offers generous health, dental, and vision benefits as well as 401(k) company match.

¿Listo para postularte en Inferact?
Postúlate en Inferact

Cómo se compara este salario de SRE

Este puesto paga $300,000/yrpor encima de el rango típico para los puestos de SRE.

$180,000 la mediana de $217,000 $310,000

Rango típico $192,500–$269,750/yr, a partir de 52 ofertas comparables de SRE en JobsRadar (salario anualizado en USD). Ver datos salariales de SRE →

Empleos similares

Specter
Platform Site Reliability Engineer
Specter
⚡ Postúlate pronto San Francisco Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2h
Anthropic
Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic
⚡ Postúlate pronto Remote-Friendly (Travel-Requir... · restringido por ubicación $405,000–$485,000
● Nuevo 👁 Visto ✓ Postulado hace 5h
Anthropic
Staff Software Engineer, AI Reliability
Anthropic
⚡ Postúlate pronto San Francisco, CA | New York C... Presencial $325,000–$485,000
● Nuevo 👁 Visto ✓ Postulado hace 5h
Waabi
Vehicle Reliability Engineer
Waabi
⚡ Postúlate pronto Dallas, TX Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 19h
Pinterest
Sr. Site Reliability Engineer, tvScientific
Pinterest
⚡ Postúlate pronto San Francisco, CA, US; Remote,... · restringido por ubicación $139,764–$287,749
● Nuevo 👁 Visto ✓ Postulado hace 20h
Pinterest
Site Reliability Engineer II, tvScientific
Pinterest
⚡ Postúlate pronto San Francisco, CA, US; Remote,... · restringido por ubicación $114,297–$235,319
● Nuevo 👁 Visto ✓ Postulado hace 20h
Ōura
Hardware Reliability Engineer
Ōura
⚡ Postúlate pronto Hybrid - San Francisco, Califo... Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 1d
Sierra
Software Engineer, Site Reliability (SRE)
Sierra
⚡ Postúlate pronto San Francisco, CA Presencial $230,000–$390,000
● Nuevo 👁 Visto ✓ Postulado hace 1d
Okta
Manager, Site Reliability Engineering
Okta
⚡ Postúlate pronto Bellevue, Washington; Chicago,... Presencial $204,000–$306,000
● Nuevo 👁 Visto ✓ Postulado hace 1d

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Inferact

Ver todos los empleos en Inferact →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis