About this Machine Learning Engineer role at Attain
About Attain
Built for consumers and companies, alike.
Klover's engineering team powers one of the fastest-growing fintech platforms in the U.S., supporting over one million active users each month. Our systems process and move more than $1.5 billion annually, enabling real-time access to financial tools, rewards, and services that help people improve their day-to-day lives.
As part of this team, you'll help design, build, and scale the systems that underpin Klover's core products and platform. You'll work on high-impact, production-grade systems that prioritize reliability, security, and performance, and that integrate with a broad ecosystem of internal and external services. The work you do will directly shape how users interact with Klover's products, access their money, and experience transparent, low-fee financial services.
Klover engineers collaborate closely with colleagues across backend, frontend, data science, and product teams to deliver scalable, high-quality solutions for a rapidly growing user base. You'll have the opportunity to work with modern technologies and architectures while helping define and evolve the next generation of inclusive, data-powered financial products—building systems and interfaces that emphasize reliability, privacy, and performance at scale.
About the role
Attain is seeking a Senior/Staff Machine Learning Engineer to own our production ML systems and build out the MLOps platform infrastructure that powers our suite of B2C financial services. This role will be highly hands-on and infrastructure-first, focused on designing, building, and operating the pipelines, platforms, and tooling that take models from experiment to reliable production service across our app portfolio—and on keeping those systems healthy, performant, and cost-effective once they're live.
You will work on the systems and infrastructure behind our high-impact predictive models, including the pipelines, feature infrastructure, model-serving, CI/CD, and observability that keep them reproducible, automated, monitored, and fast in production. Day to day, this means building the platform and automation that let us move fast without sacrificing performance—streamlining retraining and rollouts, tuning systems for speed and efficiency, and building the metrics and alerting that give us confidence to ship—while enabling data scientists to deploy and iterate on models quickly and safely. The ideal candidate combines strong software and platform engineering fundamentals with practical MLOps experience building and operating production ML systems from scratch, and treats modern AI tooling as a first-class part of how the work gets done—directing coding agents to write, test, and ship infrastructure code, with the judgment to know when to verify their work.
Attain Office Hybrid Schedule:
- Chicago, IL: 4 days in-office; 1 day remote
What a typical week might look like
- Build, deploy, and operate the production ML systems at the core of our EWA product, with a focus on reliability, performance, and fast, high-quality execution
- Build and improve the pipelines and serving infrastructure behind our predictive models across consumer decisioning, fraud, churn, transaction intelligence, and other business-critical use cases
- Own the production side of the model lifecycle: feature pipelines, deployment, CI/CD, monitoring, and automated retraining
- Build and maintain reusable modeling pipelines, feature engineering systems, model-serving infrastructure, and production-quality code, deployed via Terraform and CI/CD into our GCP + Kubernetes environment
- Instrument models and pipelines with monitoring, alerting, and automated retraining—defining the metrics and dashboards (e.g., Prometheus/Grafana) that surface drift and degradation and give us confidence to ship
- Direct AI coding agents as a force multiplier to write, test, and ship infrastructure and pipeline code—and apply strong judgment about when to trust their output and when to verify it yourself
- Automate manual, repetitive steps in the ML lifecycle so the team can move faster without sacrificing reliability
- Partner with data scientists to give them fast, safe paths to deploy, iterate on, and retrain models in production
- Collaborate with analysts, platform engineers, product managers, and business stakeholders to deliver ML systems with quality, efficiency, and precision
- Identify new areas where platform improvements, automation, and MLOps tooling can improve product velocity and business outcomes
Preferred Qualifications
- 5+ years of direct experience as a Machine Learning Engineer, ML Platform Engineer, MLOps Engineer, Applied Scientist or similar role building and operating production ML systems
- Strongly preferred: degree in STEM field such as Computer Science, Statistics, Economics, Mathematics, Engineering, Physics, Operations Research, or a related quantitative field
- Demonstrated ability to apply critical thinking, abstract reasoning, and sound engineering judgment to complex, ambiguous technical and business problems
- Strong expertise deploying, serving, monitoring, and operating ML models in production—including feature engineering systems, training/serving parity, retraining, and model performance diagnostics
- Experience building low-latency online model serving (e.g., gRPC/microservices, ideally with a service mesh such as Istio) for real-time decisioning
- Hands-on MLOps experience: pipelines, CI/CD for ML, containerization (Docker), orchestration (Kubernetes), infrastructure-as-code (e.g., Terraform), and workflow schedulers (e.g., Airflow)
- Experience with model versioning, reproducibility, and safe progressive rollout (shadow, canary, champion-challenger) of models in production
- Demonstrated fluency directing AI coding agents (e.g., Claude Code, Cursor, or similar) to build, operate, and debug real ML systems—with experienced judgment on verifying their work
- A track record of replacing manual, repetitive ML workflows with durable automation
- Experience building the infrastructure behind high-impact applied ML use cases such as credit decisioning, risk modeling, fraud, churn, or consumer behavior modeling
- Familiarity with model explainability, auditability, and the compliance considerations of regulated decisioning (a plus for credit/fintech contexts)
- Strong software and platform engineering fundamentals
- Strong Python coding skills, with the ability to build pipelines, services, and production-quality tooling from scratch; experience with a systems or backend language such as Go or Rust is a plus
- Experience with distributed computing and GPU-accelerated workloads (e.g., Spark, Ray, Dask, or distributed training/inference), including scaling data and model pipelines across clusters
- Strong SQL skills and experience with cloud data warehouses and operational databases (e.g., BigQuery, Spanner), including working with large, messy, real-world datasets
- Experience with observability tools such as Prometheus, Grafana, or Datadog
- Experience with cloud computing services or platforms; GCP preferred
- Willingness to roll up your sleeves and wear multiple hats across engineering, infrastructure, and ML execution based on business needs
- Strong written and verbal communication skills, including the ability to explain technical topics to both technical and non-technical audiences
We are excited to hear from you.
At Attain, we are passionate about finding people to continuously help us grow our organization. We encourage you to apply, even if your experience doesn’t match every detail on the job description. If we don’t see something that immediately fits, we will keep your resume on file for future opportunities.