À propos de ce poste Machine Learning Engineer — Reinforcement Learning chez Ifm Us
The Role
We’re looking for an RL infrastructure engineer to help extend and scale our end-to-end RL training systems. You’ll work side-by-side with world-class researchers and engineers to:
- Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Integrate rollout generation, reward computation, trajectory processing, policy updates, and weight synchronization
- Build robust config + launch systems across multi-node, multi-GPU clusters
- Own experiment tracking, metrics logging, and job monitoring for external visibility
- Improve training system reliability, maintainability, and performance
Hands-on experience developing end-to-end RL infrastructure is required. Strong infrastructure and systems experience is what we value most.
Key Responsibilities
- Distributed Framework Ownership – Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures.
- Training & Inference Integration – Connect training workers with rollout generation engines (e.g., vLLM, SGLang), including trajectory exchange and policy-weight synchronization.
- Launch Config & Debugging – Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets.
- Metrics & Monitoring – Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers.
- Infra Engineering – Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
- RL Pipeline Development – Implement reward/verifier integration and trajectory processing, and validate log probabilities, token masks, and loss inputs with researchers.
- RL Execution & Recovery – Coordinate rollout and training workers, including checkpoint/restart and failure recovery; track policy versions and sample staleness when using asynchronous execution.
Qualifications
Must-Haves:
- 5+ years of experience in ML systems, infra, or distributed training
- Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
- Strong software engineering fundamentals (Python, systems design, testing)
- Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
- Ability to implement algorithms across GPUs/nodes based on mathematical specs
- Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
- Experience with large-scale machine learning workloads (strong ML fundamentals)
- Hands-on experience developing an LLM RL pipeline, with ownership across rollout/inference and distributed training integration.
- Working knowledge of policy optimization methods such as PPO or GRPO, sufficient to implement and debug sampling, log probabilities, and policy updates.
- Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
- Familiarity with performance profiling, kernel fusion, or memory optimization
- Open-source contributions or published research (MLSys, ICML, NeurIPS)
- CUDA or Triton kernel experience
- Experience with large-scale pre-training
- Experience building custom training pipelines at scale and modifying them for custom needs
- Deep familiarity with training infrastructure and performance tuning
- Experience with asynchronous RL, multi-turn or tool-using rollout environments, or custom reward/verifier systems.