Jobs › Companies › Weekday AI › Research Scientist - RLHF, RLAIF & Reward Modeling

Über diese Research Scientist - RLHF, RLAIF & Reward Modeling Stelle bei Weekday AI

Weekday AI · Vor Ort · Bengaluru, Karnataka, India

This role is for one of Weekday’s clients
Salary range: Rs 5000000 - Rs 10000000 (ie INR 50 - 100 LPA)


Min Experience: 3+ years
Location: Bengaluru, Karnataka, India
JobType: full-time

We are looking for a highly skilled and research-oriented Research Scientist with 3–6 years of experience in machine learning, reinforcement learning, and large language model (LLM) alignment. The ideal candidate will have strong hands-on experience with Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), and Reward Modeling, and will contribute to developing and improving advanced AI systems.

You will work on research problems related to model alignment, preference learning, reward optimization, evaluation, and post-training. This role requires a strong understanding of modern machine learning techniques, the ability to translate research ideas into working systems, and experience conducting rigorous experiments on large-scale models.

Requirements

Key Responsibilities

  • Design, implement, and evaluate RLHF pipelines for training and aligning large language models with human preferences.
  • Develop and improve RLAIF methodologies using AI-generated feedback, preference signals, and automated evaluation frameworks.
  • Build, train, and validate reward models that accurately capture human or AI preferences and desired model behaviors.
  • Experiment with reinforcement learning and preference optimization techniques to improve model helpfulness, accuracy, safety, and instruction following.
  • Analyze model behavior and training outcomes using quantitative evaluations, benchmarks, and controlled experiments.
  • Develop data-generation, preference-collection, ranking, and annotation strategies for alignment and post-training datasets.
  • Collaborate with research engineers and ML engineers to scale training and experimentation pipelines.
  • Investigate failure modes in reward models, preference datasets, and alignment techniques, and propose research-driven solutions.
  • Stay current with emerging research in LLM alignment, reinforcement learning, preference learning, reward modeling, and AI feedback.
  • Document experimental results and communicate research findings clearly through technical reports, presentations, and research papers.

Must-Have Skills

  • 3–6 years of hands-on experience in machine learning, deep learning, reinforcement learning, or a closely related research field.
  • Strong practical experience with RLHF (Reinforcement Learning from Human Feedback).
  • Strong understanding and hands-on experience with RLAIF (Reinforcement Learning from AI Feedback).
  • Proven experience developing, training, or evaluating reward models and preference-based learning systems.
  • Strong understanding of reinforcement learning concepts, policy optimization, reward functions, preference modeling, and model evaluation.
  • Experience working with Large Language Models (LLMs) and their training or post-training workflows.
  • Strong Python programming skills and experience with modern deep learning frameworks such as PyTorch or equivalent.
  • Ability to design experiments, interpret results, troubleshoot training issues, and derive meaningful research insights.
  • Strong mathematical and statistical foundations relevant to machine learning and reinforcement learning.

Good-to-Have Skills

  • Experience with PPO, DPO, GRPO, or other reinforcement learning and preference optimization techniques.
  • Experience working with transformer architectures and LLM fine-tuning.
  • Familiarity with distributed model training and large-scale experimentation.
  • Experience publishing research papers or contributing to open-source ML research.
  • Knowledge of model evaluation, red-teaming, AI safety, or alignment research.

Qualifications

A Master’s or Ph.D. in Computer Science, Artificial Intelligence, Machine Learning, Mathematics, Statistics, or a related technical field is preferred. Candidates with strong industry research experience and demonstrated expertise in RLHF, RLAIF, and reward modeling are encouraged to apply.

Bereit, sich bei Weekday AI zu bewerben?
Bei Weekday AI bewerben

Über Weekday AI

At Weekday (backed by YC; also Product Hunt #1 product of the day), we are building the next frontier in hiring. We have built the largest database of white collar talent in India and have built outreach tools on top of it to generate highest response rates.

Alle Jobs bei Weekday AI ansehen →

Ähnliche Jobs

Quince
Senior Data Scientist ( Data Scientist III : Supply Chain Operations Research)
Quince
⚡ Früh bewerben Bengaluru, Karnataka, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 4 Tg.
HP
Applied Research - Data Scientist
HP
⚡ Früh bewerben Bengaluru, Karnataka, India Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 3 Wo.
Percepta
Research Scientist – World Modeling
Percepta
⚡ Früh bewerben New York City Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Percepta
Research Scientist – Optimization
Percepta
⚡ Früh bewerben New York City Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Percepta
Research Scientist – Reinforcement Learning (RL)
Percepta
⚡ Früh bewerben New York City Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 1 Std.
Nahc
Research Scientist
Nahc
⚡ Früh bewerben Hong Kong Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 5 Std.
SA
Machine Learning Research Scientist, Post-Training
Scale AI
⚡ Früh bewerben San Francisco, CA; Seattle, WA... Vor Ort $165,600–$207,000
● Neu 👁 Gesehen ✓ Beworben vor 9 Std.
SA
Research Scientist, Agent Robustness
Scale AI
⚡ Früh bewerben San Francisco, CA; New York, N... Vor Ort $216,000–$270,000
● Neu 👁 Gesehen ✓ Beworben vor 9 Std.
SA
Senior / Staff Machine Learning Research Scientist, Agents
Scale AI
⚡ Früh bewerben San Francisco, CA; Seattle, WA... Vor Ort $302,400–$378,000
● Neu 👁 Gesehen ✓ Beworben vor 9 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Weekday AI

Alle Jobs bei Weekday AI ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos