Jobs Companies Epsilon Labs, Inc. Software Engineer - ML Infrastructure

Über diese Software Engineer - ML Infrastructure Stelle bei Epsilon Labs, Inc.

Epsilon Labs, Inc. · Vor Ort · San Francisco, CA

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

 

Role Overview

We're seeking a Software Engineer to build the ML infrastructure and data systems that let our research team train and ship state-of-the-art models for medical imaging. Sitting in the Engineering team and working closely with research, you'll own the data pipelines that unify live production traffic with offline datasets, the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production. This role requires someone who can move fluidly between ML systems, data engineering, distributed training, and production deployment, and who measures success by how quickly the research team can iterate.

Key Responsibilities

  • Build and optimize distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.

  • Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.

  • Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.

  • Design and implement robust data pipelines to collect, process, and store large-scale multimodal medical imaging data from both production traffic and offline sources.

  • Build centralized data storage solutions with standardized formats (e.g., protobufs) that enable efficient retrieval and training across the organization.

  • Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.

  • Contribute to production serving and deployment pipelines — model rollout, canary deployments, and monitoring — alongside the backend team.

Qualifications

  • 5+ years building ML infrastructure, data pipelines, or ML systems in production

  • Strong Python skills and expertise in PyTorch or JAX

  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient

  • Hands-on experience with data pipeline technologies (e.g., Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk) and schema design

  • Experience with distributed systems, cloud infrastructure (AWS/GCP), and containerization (Docker/Kubernetes)

  • Track record of building scalable data systems and shipping production ML infrastructure

  • Ability to move quickly and handle competing priorities in a fast-paced environment

Preferred Qualifications

  • Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems

  • Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production

  • Experience building internal training or experimentation platforms used by research teams

  • Experience supporting A/B testing and experimentation workflows, including canary deployments and monitoring statistical significance

  • Familiarity with vision-language models (VLMs) or multimodal architectures

  • Experience with medical imaging formats (DICOM) and healthcare data standards

  • Familiarity with MLOps practices and model deployment pipelines

  • Experience with privacy-preserving data systems and HIPAA compliance

Bereit, sich bei Epsilon Labs, Inc. zu bewerben?
Bei Epsilon Labs, Inc. bewerben

Ähnliche Jobs

Reflection
Member of Technical Staff - Research Software Engineer
Reflection
⚡ Früh bewerben New York, NY Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Handshake
Software Engineer, Handshake AI
Handshake
⚡ Früh bewerben San Francisco, CA Vor Ort $158,000–$175,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Crusoe
Senior Software Engineer, Streaming
Crusoe
⚡ Früh bewerben San Francisco, CA - US Vor Ort $170,000–$205,000
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Column
Software Engineer
Column
⚡ Früh bewerben San Francisco, CA Hybrid
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Epsilon Labs, Inc.
Software Engineer - Backend
Epsilon Labs, Inc.
⚡ Früh bewerben San Francisco, CA Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Redwood Materials
Infrastructure Software Engineer, Energy Storage
Redwood Materials
⚡ Früh bewerben San Francisco, California, Uni... Vor Ort $180,000–$237,500
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Redwood Materials
Embedded Software Engineer – Power Electronics, Energy Storage
Redwood Materials
⚡ Früh bewerben San Francisco, California, Uni... Vor Ort $180,000–$237,500
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Discord
Software Engineer, Distributed Systems
Discord
⚡ Früh bewerben San Francisco Bay Area Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.
Redwood Materials
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials
⚡ Früh bewerben McCarran, NV; San Francisco, C... Vor Ort $152,500–$287,500
● Neu 👁 Gesehen ✓ Beworben vor 4 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Epsilon Labs, Inc.

Alle Jobs bei Epsilon Labs, Inc. ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos