À propos de ce poste Engineer / Sr. Engineer - Data & MLOps chez Weekday AI
𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀
𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟰𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟮𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟰-𝟯𝟮 𝗟𝗣𝗔)
Experience: 2+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for a technically strong DataOps / MLOps Engineer to build, deploy, operate, and scale modern data, machine learning, and Generative AI infrastructure in enterprise cloud environments.
The role focuses on creating reliable and automated engineering pipelines across DataOps, MLOps, LLMOps, and AI platforms, with strong emphasis on Databricks and AWS. The ideal candidate will have hands-on experience with cloud data platforms, CI/CD automation, model deployment, observability, infrastructure management, and production operations.
Requirements
KEY RESPONSIBILITIES
- Design, build, and maintain scalable DataOps, MLOps, and LLMOps pipelines for data, ML, and Generative AI workloads.
- Automate data provisioning, model deployment, model evaluation, monitoring, and enterprise workflow execution.
- Implement and maintain robust CI/CD pipelines for data pipelines, ML models, containerised services, and AI applications.
- Manage, configure, optimise, and scale Databricks workspaces, clusters, jobs, and enterprise data-processing environments.
- Collaborate with Data Engineers, ML Engineers, Software Engineers, and Product teams to deploy data pipelines, feature stores, ML models, and serving systems into production.
- Implement proactive monitoring, event instrumentation, alerting, and self-healing mechanisms for data and model quality issues.
- Support incident response, production troubleshooting, infrastructure upgrades, capacity planning, and cloud resource optimisation.
- Work with DevOps, SRE, IT, and Security teams to implement governance, data lineage, compliance, security, and enterprise deployment standards.
- Monitor and troubleshoot distributed data pipelines, ETL workflows, model-serving systems, and production ML infrastructure.
- Build and maintain observability and model-monitoring capabilities using tools such as MLflow, Weights & Biases, and LangSmith.
- Develop hands-on Proofs of Concept (POCs) for modern data platforms, feature stores, streaming technologies, and ML infrastructure.
- Implement infrastructure provisioning and configuration management using tools such as Terraform, CloudFormation, or Ansible.
- Support containerised applications and ML workloads using Docker and Kubernetes.
- Participate in code reviews, engineering design discussions, on-call support, and knowledge-sharing initiatives.
- Identify opportunities to improve reliability, scalability, automation, cost efficiency, and engineering productivity.
- Stay current with evolving DataOps, MLOps, LLMOps, cloud, AI, and data-platform technologies.
- Contribute to the continuous modernisation of enterprise data and ML infrastructure.
WHAT MAKES YOU A GREAT FIT
- 2–5 years of experience across DataOps, MLOps, ML Engineering, or Data Engineering in enterprise cloud environments.
- Strong hands-on expertise in Databricks administration, workspace management, cluster optimisation, jobs, and enterprise data workloads.
- Strong experience with AWS data and ML services, including services such as SageMaker, Glue, EMR, Athena, and S3.
- Solid understanding of DevOps, DataOps, MLOps, and LLMOps methodologies and practices.
- Proven experience building CI/CD automation pipelines for containerised Python, Java, or Scala applications, microservices, and ML-serving systems.
- Strong understanding of model lifecycle management, evaluation, monitoring, governance, and observability.
- Experience with tools such as MLflow, Weights & Biases, LangSmith, or similar platforms.
- Strong understanding of databases, replication, relational and NoSQL databases, and vector databases such as Pinecone, FAISS, Milvus, or Weaviate.
- Practical experience deploying, monitoring, debugging, and supporting distributed data pipelines and ETL workflows.
- Strong Git knowledge and familiarity with standard branching and collaborative development workflows.
- Experience with Terraform, CloudFormation, Ansible, or similar infrastructure-as-code and configuration-management tools.
- Proficiency in at least one scripting/programming language such as Python, Bash, or JavaScript.
- Hands-on experience with Docker and exposure to Kubernetes orchestration.
- Good understanding of the machine learning lifecycle, feature engineering, and ML deployment workflows.
- Familiarity with PyTorch, TensorFlow, NLP, computer vision, and Generative AI concepts is desirable.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent cross-functional communication and collaboration skills.
- Demonstrated ability to promote a collaborative DevOps/DataOps/MLOps culture.
- Strong ownership mindset with the ability to support production systems and participate in on-call responsibilities.
- Ability to adapt quickly to evolving technology stacks and contribute to modernisation initiatives.