Jobs › Companies › Deutsche Bank › Site Reliability Engineer, AVP

À propos de ce poste Site Reliability Engineer, AVP chez Deutsche Bank

Deutsche Bank · Sur site · (do not use) Bangalore, Velankani Tech Park

Job Description:

Job Title: Site Reliability Engineer, AVP

Location: Bangalore, India

Corporate Title: AVP

 

Role Description

We are looking for a hands-on AVP Site Reliability Engineer who is passionate about Kubernetes platforms, deep troubleshooting, automation, and continuous improvement. You will diagnose and resolve complex production issues across application, platform, network, storage, and infrastructure layers, connecting system behavior, observability signals, and tenant impact to drive root-cause resolution.

Our SREs are thoughtful and pragmatic engineers who balance doing things right with doing what is needed now. Success in this role requires technical depth, curiosity, sound judgement, clear communication, and the ability to independently drive durable solutions across complex domains while collaborating effectively with global engineering teams.


About CaaS Private

CaaS Private is an advanced on-premises Kubernetes platform built on Google Distributed Cloud (GDC). The platform hosts critical applications. It provides a resilient, observable, and secure container platform for workloads that require enterprise-grade reliability and operational discipline.


What we’ll offer you

As part of our flexible scheme, here are just some of the benefits that you’ll enjoy

  • Best in class leave policy
  • Gender neutral parental leaves
  • 100% reimbursement under childcare assistance benefit (gender neutral)
  • Sponsorship for Industry relevant certifications and education
  • Employee Assistance Program for you and your family members
  • Comprehensive Hospitalization Insurance for you and your dependents
  • Accident and Term life Insurance
  • Complementary Health screening for 35 yrs. and above

 

Your key responsibilities

  • Define, implement, and continuously improve service-level indicators, service-level objectives, alerting standards, and error budgets for the CaaS Private platform and its critical services.
  • Build and maintain observability across metrics, logs, traces, alerts, and dashboards to provide clear insight into platform health, saturation, latency, capacity, and failure modes.
  • Investigate and resolve complex production issues across Kubernetes, Linux, networking, storage, ingress, service mesh, node services, platform dependencies, and tenant workloads.
  • Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear stakeholder communication, blameless post-incident reviews, and durable follow-up actions.
  • Automate repetitive operational tasks, diagnostics, and remediation workflows to reduce toil, improve consistency, and accelerate recovery.
  • Improve platform reliability, upgrade safety, resilience, capacity planning, performance, disaster recovery, and operational readiness.
  • Improve runbooks, dashboards, alerts, support processes, and engineering standards based on recurring issues and operational learning.
  • Partner with platform, network, security, storage, and application teams on release readiness, troubleshooting, change execution, documentation, and adoption of SRE practices.
  • Contribute to reliability-focused platform enhancements and use system-level insights to improve both platform stability and tenant experience.
  • Support the global GDC/Kubernetes platform in a 24x7 follow-the-sun model, including on-call, weekend, public-holiday rotation, and early Monday coverage as required.

 

Your skills and experience

  • A bachelor’s degree in a technical or engineering discipline, with 10–12 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps or a closely related infrastructure role.
  • Strong hands-on Kubernetes expertise, including cluster operations, upgrades, troubleshooting, networking, storage, security, Helm, Operators, and workload lifecycle management in bare-metal or private-cloud environments.
  • Demonstrated SRE mindset and experience with SLI/SLO/SLA management, error budgets, incident reduction, capacity planning, performance optimization, resilience engineering, and operational excellence.
  • Strong Linux system administration and networking fundamentals, with the ability to troubleshoot complex infrastructure and distributed-system failures.
  • Experience building and operating observability solutions using Prometheus, Grafana, Splunk, distributed tracing, logging, alerting, dashboards, and OpenTelemetry-style concepts.
  • Strong automation and infrastructure scripting experience using Python, Bash, and Ansible.
  • Experience operating and troubleshooting CI/CD and GitOps platforms such as Argo CD, Jenkins, GitHub or Bitbucket, and Artifactory.
  • Solid understanding of incident management, root-cause analysis, operational readiness, virtualization, containerization, and distributed-system behavior under failure.
  • Ability to communicate clearly, collaborate across teams, prioritise under pressure, and independently take complex problems from symptom to sustainable resolution.
  • Ability to use AI tools to improve productivity and workflows while applying critical judgement and ensuring responsible, ethical use of data and AI-generated outputs.

Skills That Will Help You Excel

  • CKA or CKAD certification, or equivalent demonstrable Kubernetes expertise.
  • Experience with Istio or Envoy, service-mesh observability, traffic management, Cilium, OPA Gatekeeper, admission controls, or policy-driven operational guardrails.
  • Experience with software-defined or container-native storage, including Ceph/Rook, CSI, or comparable storage platforms.
  • Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies.
  • Practical experience with load testing, chaos testing, failure injection, alert tuning, self-healing automation, and disaster-recovery validation.
  • Exposure to regulated or low-latency environments with strict uptime, change-control, compliance, time-synchronization, deterministic-performance, or SR-IOV workload requirements.
  • Ability to read and understand Golang code while troubleshooting platform components.

What We Value

  • Passion for troubleshooting and solving difficult technical problems.
  • Ownership, self-motivation, curiosity, and a continuous-learning mindset.
  • A pragmatic approach that balances immediate restoration with long-term engineering improvement.
  • Attention to detail, collaborative behavior, and clear written and verbal communication.
  • A focus on measurable reliability outcomes, reduced recurrence, lower operational toil, and improved tenant experience.

Technology Landscape

  • Kubernetes
  • GDC/GKE
  • Linux
  • Networking
  • Helm
  • Operators
  • Argo CD
  • GitOps
  • Prometheus
  • Grafana
  • Splunk
  • OpenTelemetry
  • OpenSearch
  • Istio
  • Ansible
  • Envoy
  • Cilium
  • Ceph/Rook
  • CSI
  • Terraform
  • Ansible
  • Python
  • Bash
  • Jenkins
  • GitHub/Bitbucket
  • Artifactory
  • SLI/SLO
  • Incident Management
  • Capacity Planning
  • Disaster Recovery
  • Security
  • Compliance

 

Proven ability to leverage AI tools to enhance productivity, optimise workflows to solve business problems, while applying critical judgment to ensure responsible and ethical use of data and AI outputs.

 

How we’ll support you

  • Training and development to help you excel in your career
  • Coaching and support from experts in your team
  • A culture of continuous learning to aid progression
  • A range of flexible benefits that you can tailor to suit your needs

 

About us and our teams

Please visit our company website for further information:

https://www.db.com/company/company.html

 

We strive for a culture in which we are empowered to excel together every day. This includes acting responsibly, thinking commercially, taking initiative and working collaboratively.

Together we share and celebrate the successes of our people. Together we are Deutsche Bank Group.

We welcome applications from all people and promote a positive, fair and inclusive work environment.

Prêt à postuler chez Deutsche Bank ?
Postuler chez Deutsche Bank

À propos de Deutsche Bank

For over 150 years, our dedication to being the Global Hausbank for our clients has been driven by our people – in around 60 countries and across more than 150 nationalities. Their deep understanding, insights, expertise, and passion help our clients navigate an increasingly complex world – be it in our Corporate Bank , our Private Bank , our Investment Bank or our Asset Management (DWS) division. Together we can make a great impact for our clients at home and abroad, securing their lasting success and financial security. More information at: Deutsche Bank Careers (db.com)

Voir tous les emplois chez Deutsche Bank →

Emplois similaires

Fronius
Data Platform Engineer - Service Reliability (m/w/d)
Fronius
⚡ Postuler tôt Thalheim bei Wels Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 30 min
Proton
Site Reliability Engineer - Observability
Proton
⚡ Postuler tôt Geneva; Paris Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Proton
Site Reliability Engineer - Storage
Proton
⚡ Postuler tôt Geneva, Switzerland, Paris, Fr... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Proton
Site Reliability Engineer - Infrastructure Systems
Proton
⚡ Postuler tôt Geneva; Paris Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 1 h
Ensono
Senior Mainframe Systems Programmer - Site Reliability Engineering
Ensono
⚡ Postuler tôt Pune, India Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
Ensono
Site Reliability Engineer
Ensono
⚡ Postuler tôt Pune, India Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Postuler tôt San Mateo, CA, United States Sur site $243,290–$295,250
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
Roblox
Senior Site Reliability Engineer, Compute
Roblox
⚡ Postuler tôt San Mateo, CA, United States Sur site $196,750–$243,290
● Nouveau 👁 Vu ✓ Postulé il y a 3 h
PwC
IN_Senior Associate_ Site Reliability Engineering_GCC_Advisory_Bangalore
PwC
⚡ Postuler tôt Bengaluru Millenia Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 h

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Deutsche Bank

Voir tous les emplois chez Deutsche Bank →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit