Jobs Companies Morgan Stanley Lead Site Reliability Engineer

À propos de ce poste Lead Site Reliability Engineer chez Morgan Stanley

Morgan Stanley · Sur site · Alpharetta, Georgia, United States of America

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime. 

Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals.


Interested in joining a team that’s eager to create, innovate and make an impact on the world? Read on.  


The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.

This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues. 

What you’ll do in the role:

Drive Reliability Engineering Practices

  • Champion SRE principles, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational risk management frameworks.

  • Define and track reliability metrics that measure platform health, customer experience, and operational effectiveness.

  • Lead initiatives focused on reducing Mean Time to Detect (MTTD), Mean Time to Identify (MTTI), and Mean Time to Restore (MTTR).

Own Production Reliability

  • Provide leadership for the proactive detection, triage, and resolution of production issues impacting business-critical applications and services.

  • Serve as the primary owner for escalated production incidents, driving resolution efforts across application, infrastructure, vendor, and external partner teams until service is restored and client impact is mitigated.

  • Establish a culture of operational excellence focused on stability, resiliency, and continuous service improvement.

  • Ensure clear, concise, and timely communication during outages, providing accurate business impact assessments and recovery updates to senior leadership.

Production Governance & Change Management

  • Serve as a key gatekeeper for the production environment, ensuring adherence to change management policies, release controls, operational readiness standards, and risk management practices.

  • Assess the operational impact of technology changes and ensure appropriate testing, rollback strategies, monitoring, and support models are in place prior to production deployment.

  • Partner with development teams throughout the software lifecycle to ensure reliability, observability, and operational supportability are built into new applications and services.

Automation & Operational Efficiency

  • Identify opportunities to eliminate manual effort and operational toil through automation, self-healing capabilities, and AI-driven operational workflows.

  • Lead the development and adoption of automation solutions that improve reliability, reduce risk, and increase operational efficiency across the organization.

  • Promote a culture of engineering-led operations and continuous process optimization.

Operational Readiness & Knowledge Management

  • Establish and maintain a comprehensive knowledge management framework, ensuring runbooks, troubleshooting guides, standards, and operational procedures are accurate, current, and accessible.

  • Drive operational readiness programs that improve first-level diagnosis and reduce dependency on development teams for routine issue resolution. Create End-to-End Know your system diagrams.

  • Ensure support teams maintain high-quality documentation and standardized troubleshooting practices to accelerate incident resolution.

Technical Leadership

  • Act as a senior technical leader and trusted advisor for reliability, resiliency, observability, and production support strategies.

  • Provide guidance on architecture reviews, platform scalability, capacity planning, disaster recovery, and resiliency testing initiatives.

  • Partner with engineering teams to identify systemic risks and implement long-term solutions that improve platform stability and customer experience.

What you’ll bring to the role:

  • 10+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end-to-end troubleshooting, networking fundamentals and appropriate attention to detail

  • BS/MS or equivalent, preferably in quantitative discipline (Computer Science, Computer Engineering).

  • 5+ years’ experience in leading a small to medium team of alike skillset.

  • 5+ years of experience in driving SRE principles and Chaos Engineering.

  • Experienced, technically hands-on professional that understands both code and infrastructure

  • Strong experience in scripting language (Shell scripting, Python, Perl, etc.) and cloud driven development

  • Strong database skills with DB2, Sybase or Oracle

  • Hands-on experience with Autosys or other batch scheduling software

  • Experience in AWS/GCP/Azure Cloud technologies

  • Working knowledge on any of the DevOps & observability tools (Grafana, Prometheus, Splunk, Kibana)

  • Solid analytical skills, problem determination, and resolution recovery processes

  • Ability to interface and cultivate excellent working relationships with technology teams, business analysts, and vendors

  • Experience in web analytics tools (preferably Adobe Experience Cloud tools) is Plus

  • Should be a fast learner of technologies in a quick paced environment.

  • Have strong organizational skills and the ability to manage multiple tasks and high-pressure situations for outage handling, management, or resolution

  • Is driven to learn about new technologies, techniques and what it takes to be an integral member of this team

  • Hands-on experience administering large-scale, high-availability systems and the tools to monitor performance and availability

  • Excellent communication and writing skills specific to technical discussions across the management layers

  • Experience with incident “on call” and ability to respond to emergencies on a 24/7 basis

  • Hands-on with AI and implementation of AI tools for operational efficiency

  • Strong ownership mentality with a focus on customer satisfaction

  • Be able to manage an outage incident, coordinating user communications, and other teams to help resolve an incident.

  • Experience working with Financial Services area will be a plus

WHAT YOU CAN EXPECT FROM MORGAN STANLEY:

At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years.  Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.

To learn more about our offices across the globe, please copy and paste https://www.morganstanley.com/about-us/global-offices​ into your browser.

Expected base pay rates for the role will be between $125,000 and $175,000 per year at the commencement of employment.  However, base pay if hired will be determined on an individualized basis and is only part of the total compensation package, which, depending on the position, may also include commission earnings, incentive compensation, discretionary bonuses, other short and long-term incentive packages, and other Morgan Stanley sponsored benefit programs.

Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background.  Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.

Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.

For more information, please visit: https://www.morganstanley.com/people-opportunities/eeo.

Prêt à postuler chez Morgan Stanley ?
Postuler chez Morgan Stanley

Comment se compare ce salaire pour SRE

Ce poste paie $150,000/yrdans la fourchette habituelle pour les postes SRE.

$96,572 la médiane $155,054 $228,754

Fourchette typique $122,087–$193,000/yr, à partir de 814 annonces SRE comparables sur JobsRadar (rémunération annualisée en USD). Voir les aperçus de salaire pour SRE →

À propos de Morgan Stanley

At Morgan Stanley, we advise, originate, trade, manage and distribute capital for people, governments and institutions, always with a standard of excellence and guided by our core values. Morgan Stanley is dedicated to providing first-class service to our clients, in a way that reflects our commitment to creating a more sustainable future and fostering stronger communities around the world. In each line of business, we strive to demonstrate our belief in the power of transformative thinking, innovative strategies and leading-edge solutions—and in the ability of capital to work for the benefit of all society. What We Do

Voir tous les emplois chez Morgan Stanley →

Emplois similaires

MS
Site Reliability Engineer
Morgan Stanley
⚡ Postuler tôt Alpharetta, Georgia, United St... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
MS
Site Reliability Engineer (SRE) - AI Platform & Cloud
Morgan Stanley
⚡ Postuler tôt Alpharetta, Georgia, United St... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 4 mois
Lloyds Banking Group
Site Reliability Engineer
Lloyds Banking Group
⚡ Postuler tôt London 1-10 Praed Mews Sur site £84,051–£93,390
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Lloyds Banking Group
Site Reliability Engineer
Lloyds Banking Group
⚡ Postuler tôt Manchester Sur site £48,987–£55,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
HP
Mechanical Reliability Engineer
HP
⚡ Postuler tôt Spring, Texas, United States o... Sur site $105,050–$161,800
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
RELX
Site Reliability Engineering Lead
RELX
⚡ Postuler tôt Florida Sur site $118,300–$219,800
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
RELX
Site Reliability Engineer II
RELX
⚡ Postuler tôt Home based-Georgia · lieu restreint $71,600–$119,400
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
RELX
FinOps Senior Site Reliability Engineer II
RELX
⚡ Postuler tôt Boca Raton, FL (Yamato) Sur site $104,900–$174,700
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
AES
Engineer, Reliability
AES
⚡ Postuler tôt US, Louisville, CO Sur site $94,000–$112,625
● Nouveau 👁 Vu ✓ Postulé il y a 4 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Morgan Stanley

Voir tous les emplois chez Morgan Stanley →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit