Jobs Companies Blue Yonder Machine Learning Engineer - Reinforcement Learning

À propos de ce poste Machine Learning Engineer - Reinforcement Learning chez Blue Yonder

Blue Yonder · Sur site · Paris

About the AI Studio

The AI Studio's mission is to find the fastest possible path to an autonomous supply chain. We're developing AI agents, learning systems, training models, and more to overcome the biggest challenges remaining in the global supply chain.

In short, we are having a lot of fun.

Your Mission In This Role

We’re looking for an ambitious ML Engineer focused on LLMs, agents, and reinforcement learning to help build the training, evaluation, and tooling systems behind robust AI decision-making products.

You’ll work across LLM fine-tuning, agent environments, reward modeling, evaluations, data pipelines, and AI workflow tooling. The role is hands-on: designing experiments, shipping production code, improving model behaviour, and building the infrastructure that lets us learn quickly from both automated and human feedback.

You’ll help shape how we use LLMs inside agentic systems, how we evaluate model and agent performance, and how we turn feedback into better training data and better behaviour.


This role requires mandatory  RL training experience with LLMs, including designing and iterating on rewards, reviewing LLM traces, identifying reward hacking or shortcut behaviour, and understanding when the reward signal, environment, or training process needs to change.

Responsibilities:

  • Design and implement LLM-powered agent environments for supply chain decision-making
  • Fine-tune, adapt, and evaluate LLMs for domain-specific reasoning and decision support
  • Design, test, and iterate on reward functions that capture the behaviors we want from LLM agents
  • Review LLM traces and rollouts to understand model reasoning, failure modes, reward hacking, and shortcut behaviour
  • Identify when an LLM is exploiting the reward function, escaping the intended RL process, or optimizing for proxy metrics instead of the real objective
  • Improve reward models, environment design, prompts, tools, and feedback loops based on observed model behaviour
  • Build evaluation frameworks to measure model quality, agent performance, robustness, and failure modes
  • Create data pipelines for training, fine-tuning, preference data, synthetic data generation, and human feedback collection
  • Develop tooling that improves how the team builds, tests, debugs, and deploys AI-assisted workflows
  • Experiment with RL, RLHF, RLAIF, reward shaping, policy optimization, and agent training techniques
  • Document what works, what fails, and why, so we can compound our learnings over time
  • Stay close to the frontier of LLMs, agents, evaluations, and applied AI engineering

We want to talk if you:

  • You've trained or fine-tuned LLMs
  • Are excited about AI-assisted tools and getting the most out of them
  • Build & customize your own AI workflows
  • Have experience working with AI agents and RL environments in production
  • Are proficient in Python and PyTorch
  • Can balance research exploration with shipping working code
  • Hands on experience with RL techniques (reward shaping, policy optimization, RLHF)
  • Thrive in fast-moving environments where priorities shift
  • Care about craft in your work
  • Are curious about why things work, not just that they work

Bonus points if:

  • You have experience with human-in-the-loop ML systems
  • You've built evaluation frameworks for open-ended tasks
  • You're familiar with supply chain, logistics, or operations domains
  • You have a side project that shows you can't stop tinkering

#LI-HG1

Our Values


If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

Prêt à postuler chez Blue Yonder ?
Postuler chez Blue Yonder

À propos de Blue Yonder

Who are we? We are a proven, passionate bunch of disruptors. Our work is all about tapping into your potential so we can deliver the best solutions and customer experiences on the planet. Collaboration, respect, and a great work-life balance earned us the title of "Best Place to Work- Employees' Choice" by Glassdoor. Our people are smart, creative, rock stars with over 400 patents and 10,000 people years of domain expertise. What do we do? Blue Yonder is the world leader in digital supply chain and omni-channel commerce fulfillment. Our intelligent, end-to-end platform enables retailers, manufacturers and logistics providers to seamlessly predict, pivot and fulfill customer demand. With Blue

Voir tous les emplois chez Blue Yonder →

Emplois similaires

Cohere
Senior ML Systems Engineer, Frameworks & Tooling
Cohere
⚡ Postuler tôt London · lieu restreint CA$250,000–CA$535,000
● Nouveau 👁 Vu ✓ Postulé il y a 5 j
Hugging Face
Senior Machine Learning Engineer, Voice Agents - EMEA Remote
Hugging Face
⚡ Postuler tôt Paris, Île-de-France, France · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 6 j
Cohere
Applied Machine Learning Engineer
Cohere
⚡ Postuler tôt London Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 1 sem.
Helsing
AI Research Engineer - ML & Signal Processing
Helsing
⚡ Postuler tôt Berlin; Munich Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
Dashlane
Machine Learning Engineer
Dashlane
⚡ Postuler tôt Paris, France Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
Alice Bob
Senior MLOps Engineer
Alice Bob
⚡ Postuler tôt Paris Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
Voodoo
Senior ML Engineer - Offline Team
Voodoo
⚡ Postuler tôt Helsinki Hybride
● Nouveau 👁 Vu ✓ Postulé il y a 2 sem.
Proton
Senior Machine Learning Engineer (SOC)
Proton
⚡ Postuler tôt Paris; Geneva Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 sem.
Arago
ML Systems Engineer — Inference Acceleration
Arago
⚡ Postuler tôt Paris Offices Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 3 sem.

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez Blue Yonder

Voir tous les emplois chez Blue Yonder →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit