รber diese Engineer / Sr. Engineer - Data & MLOps Stelle bei Weekday AI
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฐ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฏ๐ฎ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฐ-๐ฏ๐ฎ ๐๐ฃ๐)
Experience: 2+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for a technically strongย DataOps / MLOps Engineerย to build, deploy, operate, and scale modern data, machine learning, and Generative AI infrastructure in enterprise cloud environments.
The role focuses on creating reliable and automated engineering pipelines acrossย DataOps, MLOps, LLMOps, and AI platforms, with strong emphasis onย Databricks and AWS. The ideal candidate will have hands-on experience with cloud data platforms, CI/CD automation, model deployment, observability, infrastructure management, and production operations.
Requirements
KEY RESPONSIBILITIES
- Design, build, and maintain scalableย DataOps, MLOps, and LLMOps pipelinesย for data, ML, and Generative AI workloads.
- Automate data provisioning, model deployment, model evaluation, monitoring, and enterprise workflow execution.
- Implement and maintain robustย CI/CD pipelinesย for data pipelines, ML models, containerised services, and AI applications.
- Manage, configure, optimise, and scaleย Databricks workspaces, clusters, jobs, and enterprise data-processing environments.
- Collaborate with Data Engineers, ML Engineers, Software Engineers, and Product teams to deploy data pipelines, feature stores, ML models, and serving systems into production.
- Implement proactive monitoring, event instrumentation, alerting, and self-healing mechanisms for data and model quality issues.
- Support incident response, production troubleshooting, infrastructure upgrades, capacity planning, and cloud resource optimisation.
- Work with DevOps, SRE, IT, and Security teams to implement governance, data lineage, compliance, security, and enterprise deployment standards.
- Monitor and troubleshoot distributed data pipelines, ETL workflows, model-serving systems, and production ML infrastructure.
- Build and maintain observability and model-monitoring capabilities using tools such asย MLflow, Weights & Biases, and LangSmith.
- Develop hands-onย Proofs of Concept (POCs)ย for modern data platforms, feature stores, streaming technologies, and ML infrastructure.
- Implement infrastructure provisioning and configuration management using tools such asย Terraform, CloudFormation, or Ansible.
- Support containerised applications and ML workloads usingย Docker and Kubernetes.
- Participate in code reviews, engineering design discussions, on-call support, and knowledge-sharing initiatives.
- Identify opportunities to improve reliability, scalability, automation, cost efficiency, and engineering productivity.
- Stay current with evolving DataOps, MLOps, LLMOps, cloud, AI, and data-platform technologies.
- Contribute to the continuous modernisation of enterprise data and ML infrastructure.
WHAT MAKES YOU A GREAT FIT
- 2โ5 years of experienceย acrossย DataOps, MLOps, ML Engineering, or Data Engineeringย in enterprise cloud environments.
- Strong hands-on expertise inย Databricks administration, workspace management, cluster optimisation, jobs, and enterprise data workloads.
- Strong experience withย AWS data and ML services, including services such as SageMaker, Glue, EMR, Athena, and S3.
- Solid understanding ofย DevOps, DataOps, MLOps, and LLMOpsย methodologies and practices.
- Proven experience buildingย CI/CD automation pipelinesย for containerised Python, Java, or Scala applications, microservices, and ML-serving systems.
- Strong understanding of model lifecycle management, evaluation, monitoring, governance, and observability.
- Experience with tools such asย MLflow, Weights & Biases, LangSmith, or similar platforms.
- Strong understanding of databases, replication, relational and NoSQL databases, andย vector databasesย such as Pinecone, FAISS, Milvus, or Weaviate.
- Practical experience deploying, monitoring, debugging, and supporting distributed data pipelines and ETL workflows.
- Strong Git knowledge and familiarity with standard branching and collaborative development workflows.
- Experience withย Terraform, CloudFormation, Ansible, or similar infrastructure-as-code and configuration-management tools.
- Proficiency in at least one scripting/programming language such asย Python, Bash, or JavaScript.
- Hands-on experience withย Dockerย and exposure toย Kubernetesย orchestration.
- Good understanding of the machine learning lifecycle, feature engineering, and ML deployment workflows.
- Familiarity withย PyTorch, TensorFlow, NLP, computer vision, and Generative AIย concepts is desirable.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent cross-functional communication and collaboration skills.
- Demonstrated ability to promote a collaborativeย DevOps/DataOps/MLOps culture.
- Strong ownership mindset with the ability to support production systems and participate in on-call responsibilities.
- Ability to adapt quickly to evolving technology stacks and contribute to modernisation initiatives.