Jobs Companies Epsilon Labs, Inc. Software Engineer - ML Infrastructure

Über diese Software Engineer - ML Infrastructure Stelle bei Epsilon Labs, Inc.

Epsilon Labs, Inc. · Vor Ort · San Francisco, CA

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

 

Role Overview

We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.

 

Sitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.

 

Key Responsibilities

  • Partner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change

  • Build a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.

  • Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.

  • Partner with researchers to prototype new ideas and translate them into production-ready code, owning end-to-end delivery from experimentation through deployment and monitoring.

  • Contribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.

  • Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.

Qualifications

  • 6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production

  • Have 2+ years of experience building ML infrastructure or systems in production

  • Strong Python skills and expertise in PyTorch or JAX

  • Experience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research

  • Experience building infrastructure or platforms specifically for research or machine learning workflows

  • Deep experience building and operating Kubernetes and cloud infrastructure at scale

  • Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient

  • Prior experience as a technical lead or mentor for other engineers

Preferred Qualifications

  • Experience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy

  • Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems

  • Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production

  • Experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.

    • Experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.

  • Experience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows

  • Familiarity with vision-language models (VLMs) or multimodal architectures

Bereit, sich bei Epsilon Labs, Inc. zu bewerben?
Bei Epsilon Labs, Inc. bewerben

Ähnliche Jobs

Pinterest
Sr. Staff Software Engineer, Data Product Platform
Pinterest
⚡ Früh bewerben San Francisco, CA, US; Remote,... · standortgebunden $208,592–$429,454
● Neu 👁 Gesehen ✓ Beworben vor 35 Min.
Pinterest
Security Software Engineer II, Detection and Response
Pinterest
⚡ Früh bewerben San Francisco, CA, US; Remote,... · standortgebunden $123,696–$254,667
● Neu 👁 Gesehen ✓ Beworben vor 35 Min.
Pinterest
Sr. Staff Software Engineer, iOS Search and Shopping Journeys
Pinterest
⚡ Früh bewerben San Francisco, CA, US; Remote,... · standortgebunden $208,592–$429,454
● Neu 👁 Gesehen ✓ Beworben vor 35 Min.
Ryz Labs
Staff Software Engineer (Elixir)
Ryz Labs
⚡ Früh bewerben Argentina · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 34 Sek.
Conga
Staff Software Engineer, DevOps
Conga
⚡ Früh bewerben Remote United States · standortgebunden $168,675–$229,900
● Neu 👁 Gesehen ✓ Beworben vor 10 Min.
Nametag
Senior Software Engineer, Reliability
Nametag
⚡ Früh bewerben Remote · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 13 Min.
AP
Intermediate Software Developer
AppDirect
⚡ Früh bewerben Montreal, Canada Vor Ort CA$95,000–CA$120,000
● Neu 👁 Gesehen ✓ Beworben vor 17 Min.
Mitratech
Principal Software Engineer (C#/.NET)
Mitratech
⚡ Früh bewerben Remote US · standortgebunden $175,000–$185,000
● Neu 👁 Gesehen ✓ Beworben vor 22 Min.
LI
Senior Software Engineer
LINQ
⚡ Früh bewerben Denver, Colorado Vor Ort $115,000–$130,000
● Neu 👁 Gesehen ✓ Beworben vor 26 Min.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Epsilon Labs, Inc.

Alle Jobs bei Epsilon Labs, Inc. ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos