About this MLOps Engineer role at Atomic Machines
About The Role:
We are seeking an MLOps Engineer to join our AI and Modeling & Simulation org within the Data Engineering and Analytics team.
You will build and operate the infrastructure that takes AI and machine learning models from experimentation to reliable production - covering training, deployment, serving, monitoring, and continuous improvement. This is a DevOps-leaning MLOps role centered on the model feedback loop: connecting production signals and expert feedback back to training so models improve as the system operates.
We are looking for senior-level candidates who can take meaningful ownership of production ML infrastructure. The scope and seniority of the role will be shaped by the candidate's experience, technical depth, and demonstrated impact.
You will work closely with Data, AI, Process, Design, and Software engineers in a highly cross-functional environment.
What You’ll Do:
- Build and evolve the MLOps platform and CI/CD: Own the path from experiment to production, including experiment tracking, model registry, packaging, automated training and retraining, deployment, and safe rollout and rollback.
- Operate model serving infrastructure: Build reliable, scalable batch, streaming, and real-time inference for models and digital twins supporting design, process control, scheduling, and inspection.
- Build ML data and feature pipelines: Turn machine telemetry, process and knowledge graphs, images, time-series, agentic conversations, and other production data into contextualized, model-ready datasets and features.
- Maintain ML data infrastructure: Support feature-store capabilities and a lakehouse foundation using Apache Iceberg on S3, with strong data quality, lineage, versioning, and reproducibility.
- Close the model feedback loop: Build model observability and human-in-the-loop systems that capture production signals and expert corrections, version them as ground truth, and feed them into evaluation and retraining workflows.
- Create paved roads for ML development: Develop standardized tooling and workflows that enable Data and AI engineers to move quickly while maintaining production reliability and reproducibility.
- Drive technical ownership: Identify infrastructure, reliability, and scalability challenges and drive solutions from design through production. More senior candidates will have opportunities to shape architecture, technical direction, and engineering practices across the ML platform.
- Collaborate across disciplines: Work with Process, Chemical, Materials, Simulation, Software, Data, and AI engineers to define deployment, serving, and data-collection requirements.
What You’ll Need:
- 5+ years of relevant industry experience building production software, infrastructure, data, or machine learning systems. We value demonstrated technical depth, ownership, and impact over a specific number of years.
- Proven experience building and operating machine learning systems in production, with a strong MLOps/DevOps orientation.
- Strong DevOps fundamentals, including CI/CD, containers, Kubernetes, cloud infrastructure, and infrastructure-as-code.
- Proficiency in Python and SQL.
- Hands-on experience with MLflow or similar tooling for experiment tracking, model registry, and model lifecycle management.
- Experience with S3, lakehouse technologies such as Apache Iceberg, and workflow orchestration tools such as Airflow or Dagster.
- Experience building pipelines for multimodal ML data, including images, time-series, structured, and semi-structured data.
- Familiarity with manufacturing systems, sensors, process automation, or other physical-world data systems.
- Strong problem-solving skills, attention to data quality and reliability, and clear technical communication.
- Bachelor's or Master's degree in Computer Science, Data Engineering, Data Science, or a related STEM field, or equivalent practical experience.
- This role is open across multiple levels, from early in career though Staff (L4 through L6). We'll determine the appropriate level and compensation based on your experience, skills, and the scope of the role through the interview process.
Bonus Points For:
- Experience with feature stores, human-in-the-loop systems, active learning, or data-labeling infrastructure.
- Robotics or robotic automation experience, including sensors, vision systems, or robotics data.
- Experience operating ML systems in manufacturing or other physical-world environments.
- Experience building internal tools for expert feedback, labeling, model evaluation, or model interaction.
- Experience designing shared ML infrastructure or platforms used across multiple teams or applications.
The compensation for this position also includes equity and benefits.