Jobs › Companies › OpenRouter › Site Reliability Engineer, Provider Operations

À propos de ce poste Site Reliability Engineer, Provider Operations chez OpenRouter

OpenRouter · Télétravail · Remote (US)

About OpenRouter

OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.

As enterprise adoption of AI accelerates, OpenRouter sits at the center of how organizations operationalize LLMs across research, product, and production workloads.

About the Role

OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.

We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.

What You'll Do

  • Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.

  • Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.

  • Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.

  • Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.

  • Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.

  • Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.

  • Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.

About You

  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.

  • Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.

  • Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.

  • Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.

  • Calm, clear incident commander who communicates well with external partners under pressure.

  • Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.

Nice to Have

  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.

  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.

  • Background in routing, load balancing, or traffic management systems.

  • Experience with evals or synthetic monitoring for ML systems.

Prêt à postuler chez OpenRouter ?
Postuler chez OpenRouter

Emplois similaires

Prove
Site Reliability Engineer, Sr. Site Reliability Engineer
Prove
⚡ Postuler tôt United States (Remote) · lieu restreint $153,000–$171,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Hims & Hers
Sr. Reliability Engineer
Hims & Hers
⚡ Postuler tôt US Remote · lieu restreint $140,000–$165,000
● Nouveau 👁 Vu ✓ Postulé il y a 2 j
Backblaze External Website
Site Reliability Engineer III (DBA)
Backblaze External Website
⚡ Postuler tôt Remote - US · lieu restreint $125,000–$150,000
● Nouveau 👁 Vu ✓ Postulé il y a 3 j
Runlayer
Member of Technical Staff - Site Reliability
Runlayer
⚡ Postuler tôt NYC Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 4 j
Anthropic
Data Center Engineer, Reliability & Infrastructure Management – Compute Supply
Anthropic
⚡ Postuler tôt San Francisco, CA | New York C... Sur site $320,000–$405,000
● Nouveau 👁 Vu ✓ Postulé il y a 4 j
Pinterest
Sr. Site Reliability Engineer, tvScientific
Pinterest
⚡ Postuler tôt San Francisco, CA, US; Remote,... · lieu restreint $139,764–$287,749
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
Rocket Money
Senior Infrastructure Engineer, SRE
Rocket Money
⚡ Postuler tôt San Francisco, CA, Washington,... · lieu restreint $150,000–$185,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
KnowBe4
Snr. Site Reliability Engineer (Remote)
KnowBe4
⚡ Postuler tôt Remote · lieu restreint $130,000–$155,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
KnowBe4
Staff Site Reliability Engineer (Remote)
KnowBe4
⚡ Postuler tôt Remote · lieu restreint $170,000–$210,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez OpenRouter

Voir tous les emplois chez OpenRouter →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit