Über diese Engineer III, Systems Reliability Engineering Stelle bei Levi Strauss & Co.
Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger. We're a company of people who like to forge our own path and leave the world better than we found it. Who believe that what makes us different makes us stronger. So add your voice. Make an impact. Find your fit — and your future
Introduction
As an Engineer III, Systems Reliability Engineering, you will help ensure the reliability, performance, and operational excellence of our Retail Technology platforms. You will work across applications, infrastructure, and store technologies to improve service stability, automate operations, and drive continuous improvement. Reporting to the Systems Reliability Engineering leadership team, you will partner with engineers, product teams, and technology stakeholders to deliver resilient and scalable technology solutions that support our global retail business.
About the Job
- Drive reliability, availability, and performance improvements across retail technology platforms, applications, and infrastructure.
- Lead operational activities including incident response, problem management, change management, release readiness, and service transition processes.
- Develop and implement automation, observability, and monitoring solutions that improve operational efficiency and proactively identify service issues.
- Partner with product, engineering, infrastructure, security, and business teams to deliver stable, scalable, and customer-focused technology services.
- Analyze service health, capacity, performance, and reliability metrics to drive continuous improvement, operational resilience, and business continuity.
About You
- 8+ years of experience in Site Reliability Engineering (SRE), application support, infrastructure engineering, production operations, or a related technology discipline.
- Experience managing or supporting enterprise-scale applications and infrastructure in complex business environments.
- Hands-on experience with incident management, problem management, change management, and production support practices.
- Experience with observability and monitoring platforms such as New Relic, application performance monitoring (APM) tools, or similar technologies.
- Experience using automation and scripting tools such as Python, PowerShell, CI/CD platforms, or cloud technologies to improve operational efficiency and reliability.