About this Systems Operations Manager role at Wells Fargo
About this role:
Wells Fargo is seeking a...
In this role, you will:
Key Responsibilities
SRE & Reliability Engineering
- Lead and mature SRE practices across platforms and application ecosystems.
- Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Drive reliability, availability, scalability, and performance improvements.
- Establish proactive monitoring, alerting, and automated remediation strategies.
- Reduce operational toil through engineering-led solutions.
Platform & Application Support
- Own production support, platform operations, and service stability for business-critical applications.
- Ensure adherence to operational KPIs including availability, MTTR, incident reduction, and change success rates.
- Lead major incident management, problem management, and root cause analysis activities.
- Drive continuous service improvement initiatives.
Automation & Engineering Excellence
- Develop and implement automation strategies across infrastructure, application, and operational workflows.
- Automate deployment, recovery, patching, monitoring, and operational processes.
- Leverage Infrastructure as Code (IaC), CI/CD, and self-healing capabilities.
- Champion DevOps and GitOps engineering practices.
Observability & Platform Monitoring
- Establish enterprise observability capabilities across applications and platforms.
- Implement centralized logging, metrics, tracing, synthetic monitoring, and AIOps solutions.
- Drive adoption of tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, New Relic, AppDynamics, or OpenTelemetry.
- Improve visibility into platform health, customer experience, and business service performance.
Platform Transformation & Modernization
- Lead transformation initiatives involving cloud migration, platform modernization, containerization, and operational excellence.
- Partner with Architecture, Engineering, Security, and Infrastructure teams to modernize platforms.
- Drive resilience engineering, chaos testing, capacity planning, and disaster recovery improvements.
- Implement best practices for cloud-native operations and enterprise-scale support models.
Leadership & Stakeholder Management
- Build, mentor, and lead geographically distributed SRE and Support teams.
- Establish a culture of accountability, innovation, continuous learning, and operational excellence.
- Partner with business leaders, engineering teams, and senior stakeholders to align operational priorities with business objectives.
- Provide executive-level reporting on service health, reliability trends, risks, and transformation initiatives.
Required Qualifications:
- 5+ years of Systems Engineering, and Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 2+ years of Leadership experience
Desired Qualifications:
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
- 10+ years of experience in IT Operations, Production Support, SRE, Platform Engineering, or Infrastructure Operations.
- 5+ years of leadership experience managing engineering or operations teams.
- Deep understanding of Site Reliability Engineering principles and practices.
- Strong hands-on expertise in Linux/Unix, Windows, Cloud Platforms (AWS/Azure/GCP), and distributed systems.
- Experience with Kubernetes, Docker, OpenShift, or container orchestration platforms.
- Expertise in incident management, problem management, change management, and operational governance.
- Strong scripting/programming skills in Python, PowerShell, Shell, Java, or similar languages.
- Experience implementing CI/CD pipelines and Infrastructure as Code (Terraform, Ansible, CloudFormation, etc.)
- Experience in large-scale enterprise application support environments.
- Knowledge of AIOps, event correlation, automated remediation, and predictive operations.
- Experience with ServiceNow, Jira, GitHub, Azure DevOps, AutoSys, Control-M, or similar tools.
- Certifications in AWS, Azure, Kubernetes, ITIL, SRE, or DevOps.
- Experience leading cloud transformation and platform modernization programs.
Key Success Measures
- Improved Service Availability and Reliability.
- Reduction in Major Incidents and Recurring Problems.
- Improved MTTR and Incident Resolution Efficiency.
- Increased Automation Coverage and Reduced Operational Toil.
- Enhanced Observability and Early Detection Capabilities.
- Successful Delivery of Platform Modernization Initiatives.
- Improved Change Success Rate and Operational Stability.
- High Team Engagement, Capability Development, and Stakeholder Satisfaction.
Ideal Candidate Profile
A hands-on engineering leader who can operate at strategic and execution levels, drive SRE maturity, champion automation and observability, modernize platforms, and build a culture of reliability, resilience, and operational excellence across enterprise applications and infrastructure.
Job Expectations:
Posting End Date:
10 Sep 2026*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit’s risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo.
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.