Jobs Companies Flutter Entertainment Senior Reliability Engineer - FanDuel, Hybrid & Remote

Sobre este puesto de Senior Reliability Engineer - FanDuel, Hybrid & Remote en Flutter Entertainment

Flutter Entertainment · Presencial · Cluj-Napoca, Romania
Senior Reliability Engineer - FanDuel, Hybrid & Remote

Senior Platform Engineer

About Betfair Romania Development​:
Betfair Romania Development is the largest technology hub of Flutter Entertainment, with over 2,000 people powering the world’s leading sports betting and iGaming brands. Exciting, immersive and safe experiences are delivered to over 18 million customers worldwide, from our office in Cluj-Napoca. Driven by relentless innovation and commitment to excellence, we operate our own unbeatable portfolio of diverse proprietary brands such as FanDuel, PokerStars, SportsBet, Betfair, Paddy Power, or Sky Betting & Gaming.

Our Values:
The values we share at Betfair Romania Development define what makes us unique as a team. They empower us by giving meaning to our contributions, and they ensure that we consistently strive for excellence in everything we do. We are looking for passionate individuals who align with our values and are committed to making a difference.
Win together | Raise the bar | Got your back | Own it | Positive impact

About FanDuel:
FanDuel is a leading force in the sports-tech entertainment industry, redefining how fans engage with their favorite sports, teams, and leagues. As the premier gaming destination in North America, FanDuel operates across multiple verticals, including sports betting, daily fantasy sports, online gaming, advance-deposit wagering, and media.

Role Overview:

As a Senior Reliability Engineer / SRE, you’ll help improve how FanDuel’s services are built, operated, and continuously improved from a reliability perspective.

You’ll work closely with application and platform teams to understand critical customer journeys, identify reliability risks, establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs), improve operational readiness, and turn production learnings into measurable engineering improvements.

This is a hands-on engineering role. You’ll use production telemetry, incident data, service architecture, and reliability practices to understand how systems behave and where improvements are needed. You’ll help teams adopt SRE practices such as SLOs, error budgets, production readiness, incident learning, toil reduction, and operational automation, while ensuring application teams continue to own the reliability of their services.

You’ll also work closely with Observability, Performance and Resilience Engineering to connect telemetry, performance testing, Gamedays, chaos testing, and incident findings into a broader view of service reliability. Where recurring operational problems exist, you’ll help identify opportunities for automation and self-service rather than relying on repeated manual intervention.

Key Accountabilities & Responsibilities:

  • Partner with application and platform teams to understand service architecture, dependencies, critical customer journeys, and reliability risks.

  • Help teams define meaningful SLIs and SLOs that connect technical service behaviour with customer and business outcomes.

  • Support the adoption of error budgets and help teams understand how reliability performance should influence engineering priorities.

  • Assess services against reliability and production-readiness expectations, identifying gaps across observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery.

  • Use reliability maturity assessments and scorecards to help teams understand their current position and prioritise improvements with the greatest impact.

  • Support incident response and investigation for complex production issues, using logs, metrics, traces, service dependencies, and deployment information to understand system behaviour.

  • Review incidents and recurring operational issues to identify patterns, systemic risks, and opportunities for longer-term engineering improvements.

  • Improve runbooks, alerting, escalation paths, and operational practices based on production learnings.

  • Identify repetitive operational work and help design automation or self-service capabilities to reduce engineering toil.

  • Build tooling and automation that improves reliability workflows, operational readiness, investigation, and remediation.

  • Partner with Observability Engineering to ensure services have the telemetry required to measure reliability and troubleshoot production issues effectively.

  • Partner with Performance and Resilience Engineering to understand findings from performance tests, capacity assessments, Gamedays, chaos experiments, and failover testing, and translate them into actionable reliability improvements.

  • Support peak-event readiness by reviewing service health, SLOs, dependencies, operational readiness, capacity risks, and outstanding reliability findings.

  • Help application teams improve the reliability of critical customer journeys through better monitoring, dependency understanding, failure handling, and recovery practices.

  • Contribute to Reliability Engineering standards, patterns, documentation, and golden paths that can be reused across engineering teams.

  • Support the integration of reliability capabilities into CI/CD and developer workflows, including SLO-as-Code, telemetry validation, production-readiness checks, and automated operational controls.

  • Use AI-assisted investigation and automation capabilities to accelerate troubleshooting, identify recurring patterns, and reduce manual operational effort.

  • Share knowledge and support other engineers in developing stronger SRE and production-engineering practices.

Skills, Capabilities & Experience Required:

  • Strong hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering.

  • Good understanding of core SRE principles, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation.

  • Experience defining or working with SLIs and SLOs for production services.

  • Strong production troubleshooting skills, with experience investigating issues across applications, infrastructure, networks, databases, and service dependencies.

  • Good understanding of distributed systems concepts and common reliability patterns such as retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation.

  • Experience working with observability platforms such as Datadog or equivalent, with a good understanding of logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring.

  • Experience participating in incident response, post-incident reviews, and the implementation of meaningful follow-up actions.

  • Hands-on experience with Kubernetes and cloud infrastructure, preferably AWS.

  • Experience with infrastructure-as-code tools such as Terraform and modern CI/CD environments.

  • Strong automation and software engineering skills, with proficiency in at least one modern programming language such as Go, Java, Python, or JavaScript.

  • Experience identifying repetitive operational activities and replacing them with automation or self-service where appropriate.

  • Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering.

  • Ability to understand application architecture and identify reliability risks across service and infrastructure dependencies.

  • Strong analytical and problem-solving skills, with the ability to turn technical signals and production behaviour into actionable engineering improvements.

  • Good communication skills and the ability to explain reliability concepts and recommendations to engineers and engineering leadership.

  • Ability to work across multiple teams, balance competing priorities, and drive work through to measurable outcomes.

  • A mindset focused on automation, continuous improvement, knowledge sharing, and solving systemic problems rather than repeatedly addressing the same symptoms.

A Sneak Peek Into Our Tech Stack:

AWS, Kubernetes, Terraform, Helm, Ansible, Vault
Datadog, OpenTelemetry, PagerDuty
Buildkite, GitHub and infrastructure-as-code workflows
Bits AI SRE and AI-assisted investigation capabilities
Locust, AWS Resilience Hub, AWS Fault Injection Service


Benefits:

  • Hybrid & remote working options

  • €1,000 per year for self-development

  • Company share scheme

  • 25 days of annual leave per year

  • 20 days per year to work abroad

  • 5 personal days/year

  • Flexible benefits: travel, sports, hobbies

  • Extended health, dental and travel insurances

  • Customized well-being programmes

  • Career growth sessions

  • Thousands of online courses through Udemy

  • A variety of engaging office events                       



Disclaimer:

We are an inclusive employer. By embracing diverse experiences and perspectives, we create a lasting, positive impact for our employees, customers, and the communities we’re part of. You don't have to meet all the requirements listed to apply for this role. If you need any adjustments to make this role work for you, let us know, and we’ll see how we can accommodate them.


We thank all applicants for their interest; however, only the candidates who best meet the job requirements will be contacted for an interview.


By submitting your application online, you agree that your details will be used to progress your application for employment. If your application is successful, your details will be used to administer your personnel record. If your application is unsuccessful, we will retain your details for a period no longer than three years, to consider you for prospective roles within the company.

¿Listo para postularte en Flutter Entertainment?
Postúlate en Flutter Entertainment

Sobre Flutter Entertainment

Our Work Experience is the combination of everything that's unique about us: our culture, our core values, our company meetings, our commitment to sustainability, our recognition programs, but most importantly, it's our people. Our employees are self-disciplined, hard working, curious, trustworthy, humble, and truthful. They make choices according to what is best for the team, they live for opportunities to collaborate and make a difference, and they make us t he #1 Top Workplace in the area. Join us and grow your career with Flutter! Please create a candidate account after submitting your application to track it's progress.

Ver todos los empleos en Flutter Entertainment →

Empleos similares

Garmin Cluj
Site Reliability Engineer | Core Software Engineering Services Team
Garmin Cluj
⚡ Postúlate pronto Cluj-Napoca, Cluj County, Roma... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 4 meses
Garmin Cluj
Senior Site Reliability Engineer | Cloud Team
Garmin Cluj
⚡ Postúlate pronto Cluj-Napoca, Cluj County, Roma... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 5 meses
Tabs
Staff Site Reliability Engineer
Tabs
⚡ Postúlate pronto New York City, NY Presencial $220,000–$260,000
● Nuevo 👁 Visto ✓ Postulado hace 4m
Tenstorrent
Staff, Reliability Engineer
Tenstorrent
⚡ Postúlate pronto Toronto, Ontario, Canada Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 1h
Fastly
Senior SRE - Networks
Fastly
⚡ Postúlate pronto United Kingdom (Remote) Híbrido
● Nuevo 👁 Visto ✓ Postulado hace 1h
NL
Senior Cloud Platform & Site Reliability Engineering Lead
National Life Insurance Company
⚡ Postúlate pronto Addison, TX; Montpelier, VT Presencial $136,875–$200,750
● Nuevo 👁 Visto ✓ Postulado hace 2h
General Matter
DevOps / Site Reliability Engineer
General Matter
⚡ Postúlate pronto Los Angeles, CA Presencial $100,000–$200,000
● Nuevo 👁 Visto ✓ Postulado hace 2h
Hermeus
Senior Build Reliability Engineer
Hermeus
⚡ Postúlate pronto Los Angeles, CA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 3h
Zscaler
Staff Site Reliability Engineer (Production Engineer)- Federal
Zscaler
⚡ Postúlate pronto Bellevue, Washington, USA; Bos... Híbrido $119,000–$170,000
● Nuevo 👁 Visto ✓ Postulado hace 3h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Flutter Entertainment

Ver todos los empleos en Flutter Entertainment →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis