Jobs Companies Sporty Group Weekend Site Reliability Engineer

Über diese Weekend Site Reliability Engineer Stelle bei Sporty Group

Sporty Group · Remote · Global - Remote

What you’ll be doing

  • Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
  • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members

What you’ll bring

  • 3+ years DevOps / platform engineering experience
  • Must be based in Europe or Asia or LatAM
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
  • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
  • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous

Our stack

  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch

What’s in it for you

  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities

Interview process

  • Remote video screening with our Talent Acquisition Team 
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)

If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.


#LI-remote

Bereit, sich bei Sporty Group zu bewerben?
Bei Sporty Group bewerben

Ähnliche Jobs

Yuno
Site Reliability Engineer
Yuno
⚡ Früh bewerben Europe · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 4 Wo.
MUFG
Senior Middleware Ops & SRE Engineer
MUFG
⚡ Früh bewerben MUFG Global Service Private Lt... · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
Tulip Interfaces
Senior Site Reliability Engineer
Tulip Interfaces
⚡ Früh bewerben Somerville, MA Hybrid $160,000–$200,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Mon.
NEORIS
Senior Site Reliability Engineer
NEORIS
⚡ Früh bewerben Colombia Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 3 Mon.
Tailor
SRE (Site Reliability Engineer)
Tailor
⚡ Früh bewerben Tokyo · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Okta
Staff Site Reliability Engineer - Kubernetes
Okta
⚡ Früh bewerben Bellevue, Washington; Chicago,... Vor Ort $194,000–$267,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Okta
SRE Operations Engineer
Okta
⚡ Früh bewerben Bengaluru, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Okta
Senior Database Reliability Engineer (DBRE)
Okta
⚡ Früh bewerben Bellevue, Washington; Chicago,... Vor Ort $160,000–$220,000
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Cardinal Health
Engineer, IT Infrastructure and System Administration - IBMi Engineer
Cardinal Health
⚡ Früh bewerben Philippines-Bonifacio Global C... · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 5 Tg.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Sporty Group

Alle Jobs bei Sporty Group ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos