Jobs Companies Morgan Stanley Site Reliability Engineer (SRE) - AI Platform & Cloud

Sobre este puesto de Site Reliability Engineer (SRE) - AI Platform & Cloud en Morgan Stanley

Morgan Stanley · Presencial · Alpharetta, Georgia, United States of America

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.

This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.

Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.

Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm’s Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses. 

This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.

As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.

The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements. 

What you'll do in the role:

  • Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)

  • Design and build automation for core platform capabilities, reducing manual toil

  • Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.

  • Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards

  • Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation

  • Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting

  • Optimize cost vs. performance tradeoffs in large-scale compute environments

  • Harden systems for security, compliance, auditability, and data governance

  • Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems

  • Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms

  • Maintain runbooks, operational playbooks, documentation, and training materials

  • Participate in on-call rotations and respond to production incidents 24/7 as needed

  • Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability

What you'll bring to the role:

  • Bachelor’s or Master’s degree in Computer Science or related field, or equivalent job experience 

  • 5 years of production experience in SRE / Infrastructure / ops for large-scale systems

  • Strong programming/scripting skills (Python, Go, Java, or equivalent)

  • Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)

  • Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)

  • Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures

  • Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)

  • Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)

  • Solid experience in capacity planning, performance tuning, scaling, and incident response

  • Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements

  • Experience in regulated environments (financial services, compliance, audit, security) is a strong plus

  • Excellent communication, documentation, and cross-team collaboration skills

  • Proven track record of reducing operational toil via automation

Nice to have

  • Understanding of SRE techniques.

  • Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.

  • Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.

  • Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)

  • Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.

  • Experience working with Generative AI development, embeddings, fine tuning of Generative AI models. 

  • Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)

  • Understanding of ModelOps/ ML Ops/ LLM Op.

  • Experience with chaos engineering, canary deployments, blue/green rollouts

We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.

All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren’t just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.

Wherever you are in our 1,200 global offices, you’ll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.

At Morgan Stanley Alpharetta, we support the Firm’s global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.

Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.

It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.

Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).

WHAT YOU CAN EXPECT FROM MORGAN STANLEY:

At Morgan Stanley, we raise, manage and allocate capital for our clients – helping them reach their goals. We do it in a way that’s differentiated – and we’ve done that for 90 years.  Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you’ll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There’s also ample opportunity to move about the business for those who show passion and grit in their work.

To learn more about our offices across the globe, please copy and paste https://www.morganstanley.com/about-us/global-offices​ into your browser.

Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background.  Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.

Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.

For more information, please visit: https://www.morganstanley.com/people-opportunities/eeo.

¿Listo para postularte en Morgan Stanley?
Postúlate en Morgan Stanley

Sobre Morgan Stanley

At Morgan Stanley, we advise, originate, trade, manage and distribute capital for people, governments and institutions, always with a standard of excellence and guided by our core values. Morgan Stanley is dedicated to providing first-class service to our clients, in a way that reflects our commitment to creating a more sustainable future and fostering stronger communities around the world. In each line of business, we strive to demonstrate our belief in the power of transformative thinking, innovative strategies and leading-edge solutions—and in the ability of capital to work for the benefit of all society. What We Do

Ver todos los empleos en Morgan Stanley →

Empleos similares

MS
Lead Site Reliability Engineer
Morgan Stanley
⚡ Postúlate pronto Alpharetta, Georgia, United St... Presencial $125,000–$175,000
● Nuevo 👁 Visto ✓ Postulado hace 1sem
MS
Site Reliability Engineer
Morgan Stanley
⚡ Postúlate pronto Alpharetta, Georgia, United St... Presencial
● Nuevo 👁 Visto ✓ Postulado hace 2sem
Lloyds Banking Group
Site Reliability Engineer
Lloyds Banking Group
⚡ Postúlate pronto London 1-10 Praed Mews Presencial £84,051–£93,390
● Nuevo 👁 Visto ✓ Postulado hace 4h
Lloyds Banking Group
Site Reliability Engineer
Lloyds Banking Group
⚡ Postúlate pronto Manchester Presencial £48,987–£55,000
● Nuevo 👁 Visto ✓ Postulado hace 4h
HP
Mechanical Reliability Engineer
HP
⚡ Postúlate pronto Spring, Texas, United States o... Presencial $105,050–$161,800
● Nuevo 👁 Visto ✓ Postulado hace 4h
RELX
Site Reliability Engineering Lead
RELX
⚡ Postúlate pronto Florida Presencial $118,300–$219,800
● Nuevo 👁 Visto ✓ Postulado hace 6h
RELX
Site Reliability Engineer II
RELX
⚡ Postúlate pronto Home based-Georgia · restringido por ubicación $71,600–$119,400
● Nuevo 👁 Visto ✓ Postulado hace 6h
RELX
FinOps Senior Site Reliability Engineer II
RELX
⚡ Postúlate pronto Boca Raton, FL (Yamato) Presencial $104,900–$174,700
● Nuevo 👁 Visto ✓ Postulado hace 6h
AES
Engineer, Reliability
AES
⚡ Postúlate pronto US, Louisville, CO Presencial $94,000–$112,625
● Nuevo 👁 Visto ✓ Postulado hace 7h

Regístrate para recibir sugerencias adaptadas a los empleos que abres y las búsquedas que guardas.

Más empleos en Morgan Stanley

Ver todos los empleos en Morgan Stanley →

Postúlate ahora
🤖

Un momento — para

JobsRadar se creó para personas reales que están pasando un mal momento en su búsqueda de empleo — no para solicitudes automatizadas. Estás haciendo clic demasiado rápido y ahora estás bloqueado temporalmente.

Vuelve más tarde. Si de verdad estás buscando empleo, cuentas con nosotros — solo compórtate como una persona.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Toma ventaja en tu búsqueda de empleo.

Únete a nuestro canal de Telegram para lo que te ayuda a conseguir el puesto — referencias salariales, el pulso semanal del mercado y avisos de nuevas funciones. Sin spam, solo señal.

Únete al canal — es gratis