About this DevOps Engineer role at Weekday AI
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฎ๐ฏ๐ด๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฎ๐ฌ๐ฒ๐ฐ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฎ.๐ฏ๐ด-๐ฎ๐ฌ.๐ฒ๐ฐ ๐๐ฃ๐)
Experience: 2+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for a hands-on and technically strongย DevOps Engineerย to build, maintain, and improve reliable, scalable, and secure cloud infrastructure and deployment environments. The role will focus onย Linux, AWS, Prometheus, and Grafana Cloud, with responsibility for infrastructure automation, monitoring, observability, deployment processes, system reliability, and production support.
The ideal candidate will have strong troubleshooting skills, a practical understanding of cloud infrastructure, and the ability to work closely with software engineering and other technical teams to improve application reliability and operational efficiency.
Requirements
KEY RESPONSIBILITIES
- Design, deploy, configure, and maintain scalableย AWS cloud infrastructureย across development, staging, and production environments.
- Administer and troubleshootย Linux-based servers and systems, including performance, availability, security, and resource utilisation.
- Support cloud services across compute, networking, storage, databases, IAM, and other AWS components.
- Implement and maintain infrastructure automation and configuration-management practices.
- Build and maintain reliableย CI/CD pipelinesย to automate application build, testing, deployment, and release processes.
- Configure and manageย Prometheusย for infrastructure and application monitoring, metrics collection, and alerting.
- Develop and maintainย Grafana Cloud dashboards, visualisations, alerts, and observability solutions.
- Monitor system health, application performance, resource utilisation, availability, and service-level indicators.
- Investigate production incidents, identify root causes, and implement permanent corrective actions.
- Troubleshoot Linux, networking, application deployment, infrastructure, and cloud-related issues.
- Improve system reliability through automation, proactive monitoring, capacity planning, and performance optimisation.
- Implement appropriate security controls across AWS infrastructure, Linux systems, access management, and deployment environments.
- Collaborate with software engineers, QA, architects, and other technical teams to improve deployment and operational processes.
- Maintain infrastructure documentation, operational runbooks, monitoring standards, and troubleshooting procedures.
- Support backup, disaster recovery, high-availability, and business-continuity requirements.
- Identify opportunities to reduce operational overhead through automation and standardisation.
- Participate in production releases, incident response, maintenance activities, and continuous improvement initiatives.
- Stay current with AWS services, DevOps practices, cloud-native technologies, observability tools, and infrastructure automation.
WHAT MAKES YOU A GREAT FIT
- 2+ years of professional experienceย in DevOps, Cloud Engineering, Site Reliability Engineering, Infrastructure Engineering, or a related role.
- Strong hands-on experience administering and troubleshootingย Linux environments.
- Good practical experience withย AWS cloud servicesย and cloud infrastructure management.
- Strong understanding of AWS compute, networking, storage, IAM, monitoring, and security concepts.
- Hands-on experience withย Prometheusย for metrics collection, monitoring, and alerting.
- Practical experience withย Grafana Cloud, including dashboards, visualisations, alerts, and observability.
- Experience building and maintainingย CI/CD pipelinesย and automated deployment workflows.
- Understanding of infrastructure-as-code and configuration-management practices.
- Good knowledge of networking fundamentals, DNS, HTTP/HTTPS, TCP/IP, load balancing, and security concepts.
- Strong troubleshooting and root-cause analysis skills across infrastructure and application environments.
- Understanding of system reliability, availability, scalability, monitoring, and performance optimisation.
- Experience with scripting or automation usingย Bash, Python, or similar technologies.
- Familiarity with Git and modern software development and deployment workflows.
- Exposure to Docker, Kubernetes, or other containerisation technologies will be an advantage.
- Strong understanding of DevOps principles, automation, observability, and production operations.
- Excellent communication and collaboration skills with the ability to work effectively with cross-functional engineering teams.
- Proactive mindset with strong ownership of infrastructure reliability, operational excellence, and continuous improvement.