Jobs Companies Plume Site Reliability Engineer (SRE) Manager

Sobre este puesto de Site Reliability Engineer (SRE) Manager en Plume

Plume · Presencial · Ljubljana, Slovenia

Life at Plume

At Plume, we believe that technology isn't about moving faster, it's about making life’s moments better. Which is why we’ve built the world's first, and only, open and hardware-independent service delivery platform for smart homes, small businesses, enterprises, and beyond. Our SaaS platform uses WiFi, advanced AI, and machine learning to create the future of connected spaces—and human experiences—at massive scale.

We now deliver services to over 60 million locations globally and have managed over 3 billion devices on our platform. We’re expanding rapidly, pioneering a new category, and we achieved our Series F funding in just four years. Our customers include many of the world's largest Internet Service Providers (ISPs) who look to Plume to help them evolve their smart home offerings while gleaning insights from their own data. 

With a bias for action and a love for being trailblazers, the team at Plume embodies a combination of relentless curiosity and imaginative innovation. We challenge ourselves to think in ways that other companies don't, work to do what should be done (rather than what can), and if we can’t do it exceptionally well, we don’t do it. It’s how we've assembled a team of world-class builders, thinkers, and doers. And it’s how we’re reinventing what’s possible every day.

 

Role Summary:

We're looking for an experienced Engineering Manager to lead our NOC L2 team. This role combines people leadership, operational ownership, and technical depth — you'll manage a team responsible for triaging, investigating, and resolving escalated network and platform incidents, while also driving process improvements that reduce incident volume and improve response times over time.

You'll be a key operational leader, responsible not just for keeping the team running day-to-day, but for building the systems, runbooks, and culture that make incident response faster, calmer, and more effective across the organization.

Responsibilities:

  • Lead, mentor, and grow a team of NOC L2 engineers responsible for escalated incident triage and resolution
  • Own on-call rotation structure, escalation policies, and incident response processes for the L2 team
  • Drive root-cause analysis and post-incident reviews, ensuring learnings translate into concrete process or system improvements
  • Partner with Engineering, Infrastructure, and Product teams to reduce recurring incident classes and improve platform reliability
  • Establish and refine runbooks, playbooks, and operational documentation to speed up incident resolution and reduce reliance on tribal knowledge
  • Monitor and report on key operational metrics (MTTD, MTTR, incident volume, escalation rates) to leadership
  • Manage team schedules, coverage, and on-call rotations to ensure 24/7 operational readiness
  • Hire, coach, and develop engineers on the team, conducting regular 1:1s, performance reviews, and career development planning
  • Act as an escalation point for the most critical or ambiguous incidents, providing hands-on technical guidance when needed
  • Collaborate with L1 NOC leadership to ensure smooth escalation handoffs and continuous improvement of triage criteria
  • Drive a culture of blameless postmortems and continuous operational learning

Qualifications:

  • 4+ years of experience in network operations, infrastructure, or site reliability roles, with at least 2+ years in a people management or team lead capacity
  • Strong understanding of networking fundamentals (TCP/IP, DNS, routing, firewalls, VPNs) and troubleshooting methodology
  • Proven experience owning incident response processes, including on-call rotation design and escalation management
  • Experience with monitoring, alerting, and observability tools (e.g., Grafana, Prometheus, Datadog, PagerDuty, Splunk)
  • Strong track record of driving operational improvements that measurably reduce incident volume or resolution time
  • Excellent communication skills — able to translate technical incidents into clear updates for both technical and non-technical stakeholders
  • Comfortable operating in a 24/7 operational environment, including managing coverage across shifts and time zones

Nice-to-Haves:

  • Experience in telecom, ISP, networking hardware, or connected-device industries
  • Familiarity with cloud infrastructure (AWS, GCP, or Azure) and container orchestration (Kubernetes, Docker)
  • Experience with automation/scripting to reduce manual operational toil (Python, Bash, or similar)
  • Prior experience scaling a NOC or SRE team through significant growth
  • ITIL or similar operational framework certification/experience

About Plume

As the creator of the only open, hardware-independent, cloud-controlled experience platform for ISPs and their subscribers, Plume partners with over 400 ISP customers, including some of the world’s largest such as Charter, Liberty Global, and J:COM. 

Using OpenSync, the most widely supported open-source, silicon-to-cloud framework for smart spaces, Plume’s software-defined network allows ISPs to decouple their service offerings from hardware and rapidly curate and deliver new services over a multi-vendor, open-platform architecture.  

Plume is an equal opportunity workplace that maintains a continuing policy of nondiscrimination in all employment practices and decisions, ensuring equal employment opportunities for all qualified individuals without regard to race, color, creed, religion, sex, national origin, age, physical or mental disability, sexual orientation, gender identity, marital status, pregnancy, childbirth or related individual conditions, medical conditions (as defined by state law), military or veteran status, or any other characteristic protected by federal, state or local law.

¿Listo para postularte en Plume?
Postúlate en Plume

Sobre Plume

 

 

Ver todos los empleos en Plume →

Empleos similares

Planet
Senior Site Reliability Engineer
Planet
⚡ Postúlate pronto Berlin, Germany; Haarlem, Neth... Híbrido €77,000–€96,300
● Nuevo 👁 Visto ✓ Postulado hace 11h
Pragmatike
Senior Site Reliability Engineer / Kubernetes (Remote)
Pragmatike
⚡ Postúlate pronto Italy · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 17m
Pragmatike
Senior Site Reliability Engineer / Kubernetes (Remote)
Pragmatike
⚡ Postúlate pronto Albania · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 17m
Pragmatike
Senior Site Reliability Engineer / Kubernetes (Remote)
Pragmatike
⚡ Postúlate pronto Greece · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 17m
AX
Senior Site Reliability Engineer I
Axon
⚡ Postúlate pronto Seattle, Washington, United St... Híbrido $134,250–$214,800
● Nuevo 👁 Visto ✓ Postulado hace 1h
AX
Site Reliability Engineer (Australia, Remote)
Axon
⚡ Postúlate pronto Australia Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1h
AX
Sr. Site Reliability Engineer I
Axon
⚡ Postúlate pronto Seattle, Washington, United St... Híbrido $134,250–$214,800
● Nuevo 👁 Visto ✓ Postulado hace 1h
Robinhood
Senior Cloud Engineer
Robinhood
⚡ Postúlate pronto London, UK Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2d
Speechify
Software Engineer, Data Infrastructure & Acquisition - Ljubljana, Slovenia
Speechify
⚡ Postúlate pronto Ljubljana, Slovenia Presencial
● Nuevo 👁 Visto ✓ Postulado hace 1sem

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Plume

Ver todos los empleos en Plume →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis