À propos de ce poste Staff Engineer chez Levi Strauss & Co.
Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger. We're a company of people who like to forge our own path and leave the world better than we found it. Who believe that what makes us different makes us stronger. So add your voice. Make an impact. Find your fit — and your future.
Introduction
As a Staff Production Engineer, Reliability Enablement, you will partner with consumer experience engineering teams to elevate reliability, observability, incident analysis, and operational readiness across our digital platforms. You will work across domains including web, mobile, commerce, identity, loyalty, and payments to build reusable capabilities, coach engineers, and establish organization-wide reliability standards. This role reports to the Senior Director, Consumer Experiences Engineering.
About the Job
- Define and drive operational readiness standards, including observability, alerting, service level objectives, runbooks, and production readiness reviews across consumer-facing domains.
- Build reusable reliability capabilities such as reference implementations, instrumentation frameworks, synthetic monitoring solutions, load testing harnesses, and cross-domain dashboards.
- Establish and coach teams on incident analysis methodologies, helping engineers identify root causes, systemic failures, and long-term corrective actions.
- Identify cross-domain reliability risks and recurring failure patterns, facilitating investigations and coordinating improvements across multiple engineering teams.
- Partner with Platform Engineering and engineering leaders to improve reliability maturity, increase team self-sufficiency, and drive continuous operational improvement.
About You
- 10+ years of experience building, operating, and supporting production software systems, including distributed systems and service integrations.
- Demonstrated experience leading production incidents, conducting root cause analyses, and delivering long-term reliability improvements.
- Hands-on proficiency in one or more modern programming languages such as Java, Node.js, TypeScript, or Go, with the ability to develop tooling and reference implementations.
- Deep expertise in observability practices, including metrics, distributed tracing, structured logging, service level indicators/objectives, error budgets, and alerting strategies.
- Experience improving engineering team capabilities through mentoring, coaching, and influencing cross-functional teams without direct authority.