Jobs Companies Coupang Sr. Director, Back-End Engineering

Sobre este puesto de Sr. Director, Back-End Engineering en Coupang

Coupang · Presencial · Seoul, South Korea

Sr. Director, Site Reliability Engineering

Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.

Key Responsibilities

  • Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.
  • Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.
  • Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.
  • Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.
  • Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.
  • Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.
  • Build and scale a world-class SRE organization capable of influencing engineering practices across the company.

Reliability Strategy, SLOs & Engineering Governance

  • Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.
  • Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.
  • Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.
  • Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.
  • Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.

Incident Management & Autonomous Operations

  • Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.
  • Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.
  • Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.
  • Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.
  • Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.

Disaster Recovery, Resilience & Capacity

  • Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.
  • Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.
  • Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.
  • Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.
  • Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.

Observability, Testing & Reliability Intelligence

  • Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.
  • Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.
  • Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.
  • Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.

Talent Leadership & Organization

  • Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors.
  • Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench.
  • Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations.
  • Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective.
  • Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering.

Technical Leadership & Architecture

  • Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure.
  • Define resilient patterns for redundancy, failover, traffic management, data recovery, workload prioritization, and dependency isolation.
  • Guide architecture reviews and platform standards for safe scaling, fault tolerance, and operational simplicity.
  • Balance availability, customer impact, engineering velocity, cost, and operational complexity in major technical decisions.
  • Maintain sufficient technical depth to challenge assumptions, guide principal engineers, and make high-quality decisions during critical incidents.

Execution & Impact

  • Deliver measurable improvements in availability, time to detect, time to mitigate, time to recover, incident recurrence, change-failure rate, and on-call burden.
  • Create disciplined operating rhythms, milestones, ownership models, and quarterly targets for strategic reliability programs.
  • Increase adoption of common reliability mechanisms and reduce team-specific implementations and manual operations.
  • Demonstrate business impact through improved customer experience, reduced outage exposure, stronger peak readiness, and more efficient use of infrastructure capacity.
  • Build credibility through predictable delivery, transparent risk management, and objective evidence of reliability improvement.

Essential Qualifications

  • Leadership experience in a best-in-class SRE, production engineering, infrastructure reliability, or cloud operations organization at hyperscaler, major cloud provider, global marketplace, leading fintech, or similarly scaled technology company.
  • Experience building an SRE practice comparable in maturity to leading industry organizations, rather than operating a traditional support or operations function renamed as SRE.
  • Experience with Kubernetes, service mesh, cloud-native platforms, traffic engineering, and large-scale capacity management.
  • Experience with chaos engineering, fault injection, regional resilience, and automated disaster recovery.
  • Experience building AI-assisted operations, incident intelligence, predictive reliability, or self-healing systems.
  • 15+ years of experience in software engineering, infrastructure engineering, distributed systems, or site reliability engineering.
  • 8+ years leading large-scale, multi-layer engineering organizations, including senior managers, directors, and senior individual contributors.
  • Demonstrated experience personally defining the vision and leading the architecture, build-out, launch, and scaled adoption of a company-wide SRE, reliability, resilience, or autonomous-operations program.
  • Prior experience building reliability systems and operating practices for high-scale, high-availability, customer-critical distributed systems.
  • Deep expertise in SLOs, error budgets, observability, incident management, disaster recovery, capacity planning, and resilience engineering.
  • Proven ability to lead through major incidents while also creating durable mechanisms that prevent recurrence.
  • Recognized as a visionary technology and organizational leader who can influence executive stakeholders, align multiple engineering organizations, and attract exceptional talent.
  • Proven ability to convert long-term strategy into measurable execution and company-wide adoption.
¿Listo para postularte en Coupang?
Postúlate en Coupang

Empleos similares

Coupang
[쿠팡풀필먼트서비스] 물류센터 자동화 설비 보전 관리자 (성남2센터)
Coupang
⚡ Postúlate pronto South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 4h
Coupang
Principal, Business Analyst
Coupang
⚡ Postúlate pronto Taipei, Taiwan Presencial
● Nuevo 👁 Visto ✓ Postulado hace 5h
Coupang
Sr. Director, Back-End Engineering
Coupang
⚡ Postúlate pronto Seoul, South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 7h
Coupang
[쿠팡] 카탈로그 품질 검수 및 운영 프로세스 개선 담당자 (인스펙터)
Coupang
⚡ Postúlate pronto Seoul, South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 9h
Coupang
[쿠팡풀필먼트서비스] SCM 운영지원 담당자 (OutBound Control Tower)
Coupang
⚡ Postúlate pronto Seoul, South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 12h
Coupang
Senior Product Manager (Advertiser Acquisition Product)
Coupang
⚡ Postúlate pronto Seoul, South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 12h
Coupang
[쿠팡] 시니어 운영 프로세스 개선 분석가 (Catalog Ops Improvement)
Coupang
⚡ Postúlate pronto Seoul, South Korea Presencial
● Nuevo 👁 Visto ✓ Postulado hace 15h
Coupang
Senior Executive Assistant
Coupang
⚡ Postúlate pronto Mountain View, USA Presencial $89,000–$89,000
● Nuevo 👁 Visto ✓ Postulado hace 16h
Coupang
Staff Software Engineer, Coupang Media Group
Coupang
⚡ Postúlate pronto Seattle, USA Presencial
● Nuevo 👁 Visto ✓ Postulado hace 18h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Coupang

Ver todos los empleos en Coupang →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis