À propos de ce poste Machine Learning Engineer — Pre-training (LLM) chez Ifm Us
The Role
We’re looking for a distributed ML infrastructure engineer to help extend and scale our training systems. You’ll work side-by-side with world-class researchers and engineers to:
- Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Implement distributed optimizers from mathematical specs
- Build robust config + launch systems across multi-node, multi-GPU clusters
- Own experiment tracking, metrics logging, and job monitoring for external visibility
- Improve training system reliability, maintainability, and performance
Much of the work will support large-scale pre-training, and pre-training experience is required. Strong infrastructure and systems experience is what we value most.
Key Responsibilities
- Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
- Optimizer Implementation – Translate mathematical optimizer specs into distributed implementations.
- Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets. Select and validate data, tensor, pipeline, expert, and context parallelism strategies as appropriate for the model and cluster.
- Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
- Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
- Data Loading & Checkpointing – Build and maintain data-loading and checkpoint/restart workflows, restoring model, optimizer, RNG, and data progress after interruptions.
Qualifications
Must-Haves:
- 5+ years of experience in ML systems, infra, or distributed training
- Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Strong software engineering fundamentals (Python, systems design, testing)
- Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
- Ability to implement algorithms across GPUs/nodes based on mathematical specs
- Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
- Experience with large-scale machine learning workloads (strong ML fundamentals)
- Experience with large-scale pre-training
- Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
- Familiarity with performance profiling, kernel fusion, or memory optimization
- Open-source contributions or published research (MLSys, ICML, NeurIPS)
- CUDA or Triton kernel experience
- Experience building custom training pipelines at scale and modifying them for custom needs
- Deep familiarity with training infrastructure and performance tuning