Jobs Companies Kody Senior Site Reliability Engineer

Sobre este puesto de Senior Site Reliability Engineer en Kody

Kody · Presencial · Shenzhen, Guangdong Province, China

Job Summary

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities

  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

Requirements

Qualifications & Requirements

  • Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Leadership & Operational Excellence

  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

Benefits

- Competitive package

- A dynamic and innovative team

- Collaborative, inclusive working environment where your contributions are recognized

¿Listo para postularte en Kody?
Postúlate en Kody

Sobre Kody

Kody is an exciting and fast-growing Fintech bringing the ease and optionality of online payments to brick and mortar businesses. Today we serve the hospitality industry but there's no limit to where we'll apply our technology in future.

Working closely with our customers from day 1, we've developed a new mobile point of sale platform and payments aggregator, enabling businesses to accept a multitude of payment methods through a web and app-based technology platform.

We’ve secured funding from leading names in tech investment and major hospitality chains. Kody’s visionary leadership team wants people with ambition and initiative to join us on this remarkable adventure as we scale the business and bring new products to market.

Ver todos los empleos en Kody →

Empleos similares

PlayStation Global
Software Engineer II Platform Data Reliability
PlayStation Global
⚡ Postúlate pronto United States, San Mateo, CA Presencial $150,100–$225,100
● Nuevo 👁 Visto ✓ Postulado hace 4h
PlayStation Global
Software Engineer II Data Reliability & Automation (APIs)
PlayStation Global
⚡ Postúlate pronto United States, San Diego, CA Presencial $150,100–$225,100
● Nuevo 👁 Visto ✓ Postulado hace 4h
Roku
Senior Software Engineer,  SRE
Roku
⚡ Postúlate pronto Bengaluru, India Presencial
● Nuevo 👁 Visto ✓ Postulado hace 5h
Worth AI
Senior DevOps Engineer, Infrastructure & Reliability
Worth AI
⚡ Postúlate pronto Orlando, Florida, United State... · restringido por ubicación
● Nuevo 👁 Visto ✓ Postulado hace 5h
Fivetran
Senior Staff Site Reliability Engineer
Fivetran
⚡ Postúlate pronto Oakland, California, United St... Híbrido $232,338–$290,422
● Nuevo 👁 Visto ✓ Postulado hace 6h
MongoDB
Senior Site Reliability Engineer, Fabric
MongoDB
⚡ Postúlate pronto Austin; New York City; San Fra... Presencial $127,000–$249,000
● Nuevo 👁 Visto ✓ Postulado hace 6h
Neros Technologies
Senior Test Quality & Reliability Engineer
Neros Technologies
⚡ Postúlate pronto Torrance, California, United S... Presencial $149,500–$209,000
● Nuevo 👁 Visto ✓ Postulado hace 8h
Rubrik Job Board
Software Engineer - Reliability (US Citizen)
Rubrik Job Board
⚡ Postúlate pronto Palo Alto, CA Presencial $158,000–$237,000
● Nuevo 👁 Visto ✓ Postulado hace 9h
MongoDB
Senior Site Reliability Engineer, Fleet Management
MongoDB
⚡ Postúlate pronto Austin; Boston; Chicago; Denve... Presencial $127,000–$249,000
● Nuevo 👁 Visto ✓ Postulado hace 9h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Kody

Ver todos los empleos en Kody →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis