Jobs โ€บ Companies โ€บ Weekday AI โ€บ Senior Site Reliability Engineer

รœber diese Senior Site Reliability Engineer Stelle bei Weekday AI

Weekday AI ยท Hybrid ยท Bengaluru, Karnataka, India

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿญ๐Ÿฏ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿญ๐Ÿฏ-๐Ÿฎ๐Ÿฌ ๐—Ÿ๐—ฃ๐—”)

Experience: 4+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are looking for an experiencedย Senior Site Reliability Engineer (SRE)ย to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems acrossย hybrid and multi-cloud environments.

The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience withย AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms, along with a strong understanding of production operations and distributed systems.

Requirements

Key Responsibilities

  • Define and manageย SLIs, SLOs, SLAs, error budgets, and reliability objectivesย for critical production services.
  • Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
  • Manage and support productionย Kubernetes environments, including Amazon EKS and Red Hat OpenShift.
  • Deploy and maintain containerised workloads usingย Docker, Kubernetes, and Helm.
  • Manage cloud infrastructure acrossย AWS and IBM Cloud, including hybrid-cloud environments.
  • Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
  • Develop and maintain infrastructure usingย Terraform and Infrastructure as Code (IaC)ย practices.
  • Automate operational processes, infrastructure tasks, and troubleshooting workflows usingย Python and Bash.
  • Build and enhance observability solutions usingย Prometheus, Grafana, OpenTelemetry, Thanos, and logging platforms.
  • Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
  • Participate in and leadย high-severity incident responseย and production troubleshooting.
  • Conduct root-cause analysis and lead post-incident reviews and corrective actions.
  • Develop and maintain capacity-planning and reliability-improvement strategies.
  • Implement secure, resilient, and compliant infrastructure practices across cloud environments.
  • Support disaster-recovery planning, testing, and continuous improvement.
  • Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
  • Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
  • Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.

What Makes You a Great Fit

  • 4โ€“6 years of professional experienceย in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
  • Strong hands-on experience withย AWS and Kubernetesย in production environments.
  • Experience managingย Amazon EKS, Docker, and Helm.
  • Practical experience withย Red Hat OpenShiftย is highly desirable.
  • Strong proficiency inย Terraformย and Infrastructure as Code practices.
  • Hands-on scripting and automation experience usingย Python and Bash.
  • Strong experience withย Prometheus and Grafanaย for monitoring and observability.
  • Experience withย OpenTelemetry, Thanos, logging platforms, or similar observability technologies.
  • Strong understanding ofย SLIs, SLOs, error budgets, incident management, and production troubleshooting.
  • Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
  • Experience working with hybrid or multi-cloud infrastructure, preferably includingย AWS and IBM Cloud.
  • Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
  • Experience with disaster recovery, capacity planning, and production resilience.
  • Exposure to regulated or compliance-driven environments such asย HIPAA, SOC 2, PCI DSS, or ISO 27001.
  • Strong analytical, troubleshooting, and root-cause analysis skills.
  • Excellent communication and collaboration skills.
  • Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
  • Experience mentoring engineers and contributing to technical architecture and reliability standards.
Bereit, sich bei Weekday AI zu bewerben?
Bei Weekday AI bewerben

รœber Weekday AI

At Weekday (backed by YC; also Product Hunt #1 product of the day), we are building the next frontier in hiring. We have built the largest database of white collar talent in India and have built outreach tools on top of it to generate highest response rates.

Alle Jobs bei Weekday AI ansehen โ†’

ร„hnliche Jobs

Netskope
Sr. Site Reliability Engineer, Engineering Stack Support
Netskope
โšก Frรผh bewerben Bengaluru, Karnataka, India Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 3 Tg.
Fivetran
Staff Site Reliability Engineer
Fivetran
โšก Frรผh bewerben Bengaluru, Karnataka, India, A... Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 4 Tg.
GSSTech Group
Senior DevOps / SRE Engineer
GSSTech Group
โšก Frรผh bewerben Bengaluru, Karnataka, India Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 1 Wo.
ServiceTitan
Principal Site Reliability Engineer
ServiceTitan
โšก Frรผh bewerben India Bengaluru, Karnataka Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 3 Wo.
Harness
Staff Software Engineer โ€“ AI SRE
Harness
โšก Frรผh bewerben Bengaluru, Karnataka, India Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 3 Wo.
ServiceTitan
Staff Site Reliability Engineer
ServiceTitan
โšก Frรผh bewerben India Bengaluru, Karnataka Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 3 Wo.
ServiceTitan
Senior Staff Site Reliability Engineer
ServiceTitan
โšก Frรผh bewerben India Bengaluru, Karnataka Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 4 Wo.
ZG
Lead Site Reliability Engineer
Zeta Global
โšก Frรผh bewerben Bengaluru, Karnataka, India Vor Ort
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 2 Mon.
NL
Senior Cloud Platform & Site Reliability Engineering Lead
National Life Insurance Company
โšก Frรผh bewerben Addison, TX; Montpelier, VT Vor Ort $136,875โ€“$200,750
โ— Neu ๐Ÿ‘ Gesehen โœ“ Beworben vor 4 Std.

Registrieren fรผr Vorschlรคge, die auf die von Ihnen geรถffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Weekday AI

Alle Jobs bei Weekday AI ansehen โ†’

Jetzt bewerben
๐Ÿค–

Moment โ€” langsam

JobsRadar wurde fรผr echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben โ€” nicht fรผr automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorรผbergehend blockiert.

Kommen Sie spรคter wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen โ€” verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei fรผr das, was dir hilft, die Stelle zu bekommen โ€” Gehaltsbenchmarks, den wรถchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten โ€” kostenlos