Jobs Companies Runware Senior Site Reliability Engineer

Sobre este puesto de Senior Site Reliability Engineer en Runware

Runware · Remoto · United Kingdom

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements

  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation

Bonus

  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

Benefits

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off – vacation, sick days, public holidays
  • Meaningful stock options – share in the upside you create
  • Remote-first setup – work from home anywhere we can employ you
  • Flexible hours – own your schedule outside core collaboration blocks
  • Family leave – paid maternity, paternity, and caregiver time
  • Company retreats – twice-yearly gatherings in inspiring locations

¿Listo para postularte en Runware?
Postúlate en Runware

Sobre Runware

Runware is building a leading API for all AI.

Build AI features across image, video, audio, 3D and LLMs. The lowest-cost API on the market, no infrastructure to manage. Go live in hours.

Ver todos los empleos en Runware →

Empleos similares

Carta
Senior Site Reliability Engineer
Carta
⚡ Postúlate pronto London, England, United Kingdo... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 11h
Cadillac Formula 1® Team
Enterprise Platform SRE
Cadillac Formula 1® Team
⚡ Postúlate pronto Silverstone, England, United K... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1d
Spire
Senior Software Engineer - Space Reliability
Spire
⚡ Postúlate pronto Glasgow, Scotland, United King... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1d
Scopely
Platform Engineer (Reliability) - Unannounced Project
Scopely
⚡ Postúlate pronto ES - Spain; GB - United Kingdo... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 4d
Fastly
Senior SRE - Networks
Fastly
⚡ Postúlate pronto London, United Kingdom Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 5d
Ensono
Site Reliability Engineer
Ensono
⚡ Postúlate pronto Remote - United Kingdom · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 1sem
GitLab
Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
GitLab
⚡ Postúlate pronto Remote, Canada; Remote, United... · restringido por ubicación $126,400–$314,400
● Nuevo 👁 Visto ✓ Postulado hace 1sem
Airalo
Senior Site Reliability Engineer
Airalo
⚡ Postúlate pronto Spain Remoto
● Nuevo 👁 Visto ✓ Postulado hace 1sem
ClickHouse
Senior Site Reliability Engineer- Remote
ClickHouse
⚡ Postúlate pronto United Kingdom Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1sem

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Runware

Ver todos los empleos en Runware →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis