Jobs Companies RBC Site Reliability Engineer (SRE), Cloud Operations

Sobre esta vaga de Site Reliability Engineer (SRE), Cloud Operations na RBC

RBC · Presencial · TORONTO, Ontario, Canada

Job Description

What is the opportunity?

Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms — from OpenShift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices. If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations — this is the role.

What will you do?

  • Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
  • Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
  • Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
  • Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
  • Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., ServiceNow, Prometheus) that support operational excellence.
  • Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
  • Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
  • Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.

What do you need to succeed?

Must-have

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on-call rotations, and post-incident review practices.
  • Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).

Nice-to-have

  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
  • Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
  • Experience with GPU/compute infrastructure for ML inference workloads

What’s in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • Flexible work/life balance options
  • Opportunities to do challenging work
  • Opportunities to take on progressively greater accountabilities  
  • Access to a variety of job opportunities across business

#LI-post

#TECHPJ

Job Skills

Agile Methodology, Ansible Tower, Group Problem Solving, IT System Administration, IT Systems Integration, Kubernetes, Linux, Organizational Leadership, Product Services, RedHat OpenShift Administration, Red Hat OS Administration, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems Software

Additional Job Details

Address:

RBC CENTRE, 155 WELLINGTON ST W:TORONTO

City:

Toronto

Country:

Canada

Work hours/week:

37.5

Employment Type:

Full time

Platform:

TECHNOLOGY AND OPERATIONS

Job Type:

Regular

Pay Type:

Salaried

Posted Date:

2026-07-31

Application Deadline:

2026-10-09

Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Join our Talent Community

Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.

Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com.

RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.

Pronto para se candidatar à RBC?
Candidatar-se à RBC

Sobre a RBC

Royal Bank of Canada is a global financial institution with a purpose-driven, principles-led approach to delivering leading performance. Our success comes from the 84,000+ employees who bring our vision, values and strategy to life so we can help our clients thrive and communities prosper. As Canada’s biggest bank, and one of the largest in the world based on market capitalization, we have a diversified business model with a focus on innovation and providing exceptional experiences to more than 16 million clients in Canada, the U.S. and 34 other countries. Learn more at rbc.com .‎ We are proud to support a broad range of community initiatives through donations, community investments and empl

Ver todas as vagas na RBC →

Vagas semelhantes

Okta
Staff Software Reliability Engineer - Data Platform
Okta
⚡ Candidate-se cedo Toronto, Ontario, Canada Presencial CA$160,000–CA$220,000
● Nova 👁 Vista ✓ Candidatada há 2d
Genesys
Senior Operations Reliability Engineer - IAM
Genesys
⚡ Candidate-se cedo Ontario, Canada Presencial
● Nova 👁 Vista ✓ Candidatada há 5d
Tenstorrent
Staff, Reliability Engineer
Tenstorrent
⚡ Candidate-se cedo Toronto, Ontario, Canada Híbrido
● Nova 👁 Vista ✓ Candidatada há 1sem
Funded.club
Senior Software Engineer - Site Reliability
Funded.club
⚡ Candidate-se cedo Toronto, Ontario, Canada Presencial
● Nova 👁 Vista ✓ Candidatada há 1sem
RB
Senior SRE/AIOps Engineer
RBC
⚡ Candidate-se cedo TORONTO, Ontario, Canada Presencial
● Nova 👁 Vista ✓ Candidatada há 2sem
BuildOps
Staff Software Engineer, Quality & Reliability Platform
BuildOps
⚡ Candidate-se cedo Toronto, Ontario, Canada Presencial $172,000–$229,000
● Nova 👁 Vista ✓ Candidatada há 2sem
RB
Site Reliability Engineer - SRE
RBC
⚡ Candidate-se cedo TORONTO, Ontario, Canada Presencial
● Nova 👁 Vista ✓ Candidatada há 2sem
RB
Lead Site Reliability Engineer
RBC
⚡ Candidate-se cedo TORONTO, Ontario, Canada Híbrido
● Nova 👁 Vista ✓ Candidatada há 2sem
Tipalti
Site Reliability Engineer
Tipalti
⚡ Candidate-se cedo Toronto, Ontario, Canada Presencial CA$100,000–CA$125,000
● Nova 👁 Vista ✓ Candidatada há 2sem

Cadastre-se para receber sugestões sob medida com base nas vagas que você abre e nas buscas que você salva.

Mais vagas na RBC

Ver todas as vagas na RBC →

Candidatar-se agora
🤖

Opa — calma aí

A JobsRadar foi feita para pessoas de verdade passando por um momento difícil na busca por emprego — não para requisições automatizadas. Você está clicando rápido demais e agora está temporariamente bloqueado.

Volte mais tarde. Se você está mesmo procurando emprego, estamos com você — apenas aja como um ser humano.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Ganhe vantagem na sua busca por emprego.

Entre no nosso canal do Telegram para o que ajuda você a conseguir a vaga — referências salariais, o pulso semanal do mercado e avisos de novos recursos. Sem spam, só sinal.

Entre no canal — é grátis