Über diese Senior Site Reliability Engineer Stelle bei Weekday AI
This role is for one of Weekday’s clients
Salary range: Rs 2500000 - Rs 4500000 (ie INR 25 - 45 LPA)
Min Experience: 7+ years
Location: Chennai
JobType: full-time
The Senior SRE is responsible for deployment, updates, and operational support for environments hosting our leading client’s cloud-based solutions. This role ensures operational excellence, a seamless client experience, and continuous improvement across infrastructure and delivery processes. The ideal candidate combines strong technical capabilities with the ability to lead delivery through influence and hands-on engineering expertise.
Requirements
What You’ll Do?
Operational Excellence
- Act as a senior technical authority for APAC Site Reliability Engineering activities.
- Drive best practices in reliability, operations, and engineering standards.
- Promote technical excellence, collaboration, and accountability across stakeholders.
Service Reliability & Performance
- Make infrastructure complexity transparent to both internal teams and customers, ensuring a consistently excellent client experience.
- Implement, track, and evolve service performance measures such as SLAs, SLOs, and SLIs.
- Anticipate risks related to service availability, capacity, performance regressions, and security vulnerabilities.
- Drive continuous improvement, including leading and facilitating Root Cause Analysis (RCA) activities.
Delivery & Deployment Management
- Ensure timely execution of deployments, upgrades, maintenance activities, and change requests.
- Anticipate workload, plan deliverables, and ensure qualification/validation of upcoming tasks.
- Collaborate closely with engineering to improve platform components, automation, and operational processes.
Cost, Automation & Efficiency
- Control and optimise infrastructure and operational expenditure across cloud and on-prem environments.
- Drive automation initiatives and enhance monitoring to improve reliability and reduce manual effort.
- Improve deployment workflows and operational tooling to support scalability and efficiency.
Cross-Functional Collaboration
- Communicate proactively with internal stakeholders including Pre-Sales, Project Management, CSM, and TAM teams.
- Engage with customers as needed to ensure clarity, trust, and high-quality service.
Performance Measures
- Timely delivery of planned activities
- SLA & SLO compliance
- Client satisfaction & feedback
- Operational & cloud cost optimisation
- Continuous improvement of operational practices
What We Need for You to Be Successful?
- Cloud platforms: AWS, Azure
- Containerisation & Orchestration: Kubernetes
- Infrastructure as Code: Terraform
- Configuration Management: Ansible
- Packaging & Deployment: Helm
- Databases: MariaDB, MongoDB
- Monitoring, observability, networking, and cloud security
Preferred Experience & Attributes
- Experience in complex cloud operations, SRE, or DevOps environments
- Strong problem-solving skills with a proactive mindset
- Excellent communication skills
- Ability to thrive in a fast-paced environment
Must-have skills
SRE, Helm Charts, AWS
Good-to-have skills
Python, Kubernetes