Jobs Companies Weekday AI Senior Operations Engineer

About this Senior Operations Engineer role at Weekday AI

Weekday AI · Onsite · Bengaluru, Karnataka, India

This role is for one of the Weekday's clients

Salary range: Rs 600000 - Rs 2500000 (ie INR 6 - 25 LPA)

Min Experience: 4+ years

Location: Bengaluru
JobType: full-time

We are seeking a highly skilled Senior Operations Engineer to support and optimize modern AI-powered production platforms. The ideal candidate will have strong expertise in Python, Site Reliability Engineering (SRE), DevOps, and GenAI application support. You will be responsible for ensuring the reliability, scalability, and operational excellence of AI workloads while collaborating with engineering teams to deploy and maintain production-grade Generative AI solutions.

This role is ideal for engineers who enjoy solving complex operational challenges, automating infrastructure, and supporting next-generation AI applications in cloud environments.

Requirements

Key Responsibilities

  • Manage, monitor, and optimize production environments supporting Generative AI applications and services.
  • Develop automation tools and operational utilities using Python to improve system reliability and operational efficiency.
  • Support deployment, maintenance, and troubleshooting of GenAI applications built using modern AI frameworks.
  • Monitor production systems, investigate incidents, perform root cause analysis, and implement preventive measures.
  • Collaborate with software engineering, platform, and infrastructure teams to ensure high availability and performance.
  • Deploy, maintain, and optimize AI workloads across AWS, Google Cloud Platform (GCP), or similar cloud environments.
  • Support Retrieval-Augmented Generation (RAG) pipelines and AI application workflows.
  • Implement monitoring, logging, alerting, and incident response processes for production systems.
  • Optimize Linux-based infrastructure, networking, and system performance.
  • Contribute to CI/CD automation, infrastructure improvements, and operational best practices.
  • Maintain operational documentation, runbooks, and incident management procedures.

Requirements

Must-Have Skills

  • Strong programming expertise in Python.
  • KARAT assessment completion (mandatory, if applicable to the hiring process).
  • Hands-on experience developing or supporting Generative AI (GenAI) applications.
  • Experience building or supporting applications using LangChain, LangGraph, and RAG (Retrieval-Augmented Generation) pipelines.
  • Experience working with AWS, Google Cloud Platform (GCP), or OpenAI platforms.
  • Strong experience in Site Reliability Engineering (SRE), DevOps, Production Engineering, or Production Support.
  • Good understanding of Linux systems administration and networking concepts.
  • Experience with monitoring, troubleshooting, incident management, and production operations.
  • Strong understanding of cloud-native infrastructure and operational best practices.

Good-to-Have Skills

  • Experience with LangChain.
  • Experience with LangGraph.
  • Knowledge of Kubernetes, Docker, or container orchestration platforms.
  • Familiarity with CI/CD pipelines and Infrastructure as Code (IaC).
  • Experience supporting Machine Learning or AI production workloads.
  • Exposure to observability tools, monitoring platforms, and log management solutions.
  • Knowledge of automation frameworks and scripting for infrastructure management.

Preferred Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 4–10 years of experience in SRE, DevOps, Production Engineering, or Platform Operations.
  • At least 1 year of experience supporting Generative AI or Machine Learning production workloads is preferred.
  • Experience working in cloud-native, distributed, or AI-driven environments.

Soft Skills

  • Strong analytical and troubleshooting skills.
  • Excellent problem-solving and critical thinking abilities.
  • Effective communication and cross-functional collaboration skills.
  • Ability to manage production incidents under pressure.
  • Strong ownership mindset with a focus on operational excellence.
  • Self-motivated, adaptable, and eager to learn emerging AI technologies.

Success Metrics

Success in this role will be measured by maintaining high availability, reliability, and performance of production AI platforms while minimizing system downtime and incident resolution times. Performance will also be evaluated based on the successful deployment and operational stability of GenAI applications, the effectiveness of automation initiatives, and improvements in system monitoring and observability.

Additionally, success will be reflected in proactive incident prevention, continuous optimization of cloud infrastructure, efficient support for AI workloads, and strong collaboration with engineering teams to ensure secure, scalable, and reliable production environments.

Ready to apply to Weekday AI?
Apply to Weekday AI

About Weekday AI

At Weekday (backed by YC; also Product Hunt #1 product of the day), we are building the next frontier in hiring. We have built the largest database of white collar talent in India and have built outreach tools on top of it to generate highest response rates.

See all jobs at Weekday AI →

Similar jobs

Sign up for suggestions tailored to the jobs you open and the searches you save.

More jobs at Weekday AI

See all jobs at Weekday AI →

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get an edge on your job hunt.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free