About this Lead DataOps / MLOps role at Protective
The DataOps/MLOps Lead owns the platform and operating model that lets data and ML engineers ship reliably. You build the paved paths — CI/CD, orchestration, observability, environments, and governance automation — on our Databricks Lakehouse on Microsoft Azure, so that pipelines and models move from development to governed production quickly, safely, and repeatably.
KEY RESPONSIBILITIES
• Lead CI/CD standards and pipelines in Azure DevOps (ADO) for data pipelines and ML models — build, test, and release automation, environment promotion, and repeatable, auditable deployments.
• Standardize orchestration on Dagster — reusable assets, scheduling, backfills, dependency management, and run observability across the pod's pipelines.
• Operationalize the ingestion and transformation stack — dlt (dltHub) and dbt — with automated testing, CI checks, and safe deployment of changes.
• Build MLOps foundations with the ML engineering team — MLflow model registry, Databricks Model Serving, automated deployment, monitoring, drift detection, and retraining triggers.
• Establish data and model observability — freshness, quality, lineage, latency, drift, and cost — with alerting and clear SLAs/SLOs.
• Administer and govern the Databricks Lakehouse on Azure — workspace configuration, Unity Catalog governance, access controls, and policy automation.
• Manage infrastructure as code and environments — reproducible dev/test/prod setups (e.g., Terraform), secrets management, and least-privilege access.
• Own reliability and incident practices — on-call, runbooks, root-cause analysis, and continuous improvement for data and ML services.
• Drive cost visibility and optimization (FinOps) across compute, storage, and model serving.
• Automate governance and compliance controls — audit logging, model and pipeline inventories, approval workflows, and evidence collection for a regulated environment.
• Provide technical leadership and mentoring — coach engineers on operational excellence and set the platform standards the pod builds on.
QUALIFICATIONS
• 8+ years in data, ML, or platform engineering, or in SRE/DevOps, including several years operating production data and/or ML systems.
• Demonstrated technical leadership — setting standards, building paved paths and automation, and mentoring engineers (formal people management not required, but valued).
• Strong CI/CD expertise with Azure DevOps (ADO) — build/release pipelines, environment promotion, automated testing — and Git-based workflows.
• Hands-on experience with orchestration (Dagster or equivalent) and the modern data stack — dlt (dltHub) ingestion and dbt modeling — on a Databricks lakehouse (Delta Lake).
• MLOps experience — MLflow model registry, model deployment/serving, monitoring, drift detection, and retraining automation.
• Infrastructure-as-code and cloud platform administration on Microsoft Azure (compute, storage, identity, networking basics); Terraform or equivalent.
• Strong Python and SQL for automation and tooling.
• Experience with observability/monitoring tooling and SRE practices — SLAs/SLOs, alerting, and incident management.
• Demonstrated rigor in security, access control, and secure, compliant handling of sensitive data.
• Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience.
PREFERRED QUALIFICATIONS
• Experience in financial services or insurance platform work, and familiarity with model risk and regulatory audit expectations.
• Databricks administration — Unity Catalog, cluster policies, and Mosaic AI — and familiarity with Azure Machine Learning.
• Containerization and orchestration (Docker, Kubernetes; Azure AKS or Container Apps).
• Experience with data-quality / observability tooling (e.g., Great Expectations, Monte Carlo, or similar).
• Experience automating responsible-AI and model-governance controls.
• Relevant certification such as Databricks Certified Data Engineer/ML Engineer, Microsoft Azure DevOps Engineer, or Azure Administrator.