รber diese Senior Site Reliability Engineer Stelle bei Weekday AI
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฏ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฎ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฏ-๐ฎ๐ฌ ๐๐ฃ๐)
Experience: 4+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for an experiencedย Senior Site Reliability Engineer (SRE)ย to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems acrossย hybrid and multi-cloud environments.
The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience withย AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms, along with a strong understanding of production operations and distributed systems.
Requirements
Key Responsibilities
- Define and manageย SLIs, SLOs, SLAs, error budgets, and reliability objectivesย for critical production services.
- Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
- Manage and support productionย Kubernetes environments, including Amazon EKS and Red Hat OpenShift.
- Deploy and maintain containerised workloads usingย Docker, Kubernetes, and Helm.
- Manage cloud infrastructure acrossย AWS and IBM Cloud, including hybrid-cloud environments.
- Design and maintain reliable cloud connectivity, networking, disaster-recovery, and failover solutions.
- Develop and maintain infrastructure usingย Terraform and Infrastructure as Code (IaC)ย practices.
- Automate operational processes, infrastructure tasks, and troubleshooting workflows usingย Python and Bash.
- Build and enhance observability solutions usingย Prometheus, Grafana, OpenTelemetry, Thanos, and logging platforms.
- Monitor system health, identify performance bottlenecks, and proactively address reliability risks.
- Participate in and leadย high-severity incident responseย and production troubleshooting.
- Conduct root-cause analysis and lead post-incident reviews and corrective actions.
- Develop and maintain capacity-planning and reliability-improvement strategies.
- Implement secure, resilient, and compliant infrastructure practices across cloud environments.
- Support disaster-recovery planning, testing, and continuous improvement.
- Collaborate with software engineering, platform, security, and architecture teams to improve production reliability.
- Contribute to architecture reviews, engineering standards, operational best practices, and automation initiatives.
- Mentor engineers and promote strong SRE, DevOps, observability, and production-engineering practices.
What Makes You a Great Fit
- 4โ6 years of professional experienceย in Site Reliability Engineering, DevOps, Cloud Infrastructure, or a closely related field.
- Strong hands-on experience withย AWS and Kubernetesย in production environments.
- Experience managingย Amazon EKS, Docker, and Helm.
- Practical experience withย Red Hat OpenShiftย is highly desirable.
- Strong proficiency inย Terraformย and Infrastructure as Code practices.
- Hands-on scripting and automation experience usingย Python and Bash.
- Strong experience withย Prometheus and Grafanaย for monitoring and observability.
- Experience withย OpenTelemetry, Thanos, logging platforms, or similar observability technologies.
- Strong understanding ofย SLIs, SLOs, error budgets, incident management, and production troubleshooting.
- Good understanding of DNS, TCP/IP networking, TLS, VPNs, load balancing, firewalls, and cloud connectivity.
- Experience working with hybrid or multi-cloud infrastructure, preferably includingย AWS and IBM Cloud.
- Strong understanding of containers, distributed systems, scalability, availability, and fault tolerance.
- Experience with disaster recovery, capacity planning, and production resilience.
- Exposure to regulated or compliance-driven environments such asย HIPAA, SOC 2, PCI DSS, or ISO 27001.
- Strong analytical, troubleshooting, and root-cause analysis skills.
- Excellent communication and collaboration skills.
- Ability to take ownership of critical production systems and operate effectively during high-severity incidents.
- Experience mentoring engineers and contributing to technical architecture and reliability standards.