Companies NewsBreak Research Intern, Agent RL Training

About the role

NewsBreak

About NewsBreak

Founded in 2015, NewsBreak is the Content Intelligence platform shaping the future content economy. With over 40 million monthly active users, our flagship platform delivers highly personalized local news and information powered by advanced AI, recommendation systems, and adtech.

Recognized by Fast Company as #32 on the Top Workplaces for Innovators, we're proud to be Great Place to Work® certified and home to a dynamic team of technologists, product innovators, and business leaders who are passionate about solving meaningful challenges at scale.

Together, we reached unicorn status in 2021, and we remain committed to continuing this high-growth trajectory with the right team to fulfill our mission: building the infrastructure layer for content intelligence.

If you’re inspired to dream big, innovate fast, and make a difference, we’d love to hear from you! For more information, visit www.newsbreak.com/about

About the Role

We are looking for a Research Intern to join our Agent RL Training team. You will be paired with a full-time employee as your mentor, working together to explore, from zero to one, how to apply large language models to NewsBreak’s core business, including content understanding, recommendation, agentic web browsing, and autonomous multi-step task completion.

This is a hands-on research role. You are expected to independently drive experiments, propose novel ideas, and iterate quickly. We value self-starters with deep intellectual curiosity and the drive to push boundaries in LLM post-training and agent capabilities.

Location: Onsite in Mountain View, CA office

What You’ll Work On

  • Collaborate with your full-time mentor to identify high-impact research directions for applying LLMs to NewsBreak’s products
  • Independently run end-to-end SFT experiments on LLM-based agents, and assist with RL-related exploration such as reward design and training iteration
  • Curate and build high-quality training datasets: instruction-following, preference pairs, agent trajectories, and synthetic data
  • Contribute to public publications; we encourage and support top-venue submissions during your internship

What We’re Looking For

Requirements

  • Highly motivated and committed: willing to put in extra hours when needed to push projects across the finish line
  • Genuine passion for research: you read papers for fun, tinker with models on weekends, and care deeply about advancing the field
  • Independently capable of end-to-end model SFT: with basic understanding of RL-based post-training methods (RLHF, DPO, PPO, GRPO, etc.)
  • Excellent taste in model behavior: able to reason about what “good” looks like across user-facing domains and articulate why
  • Strong Python and PyTorch skills

Preferred Qualifications

  • Publication at a top-tier venue (NeurIPS, ICML, ICLR, ACL, EMNLP, or equivalent)
  • Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM)
  • Proficiency in writing custom GPU kernels with Triton or CUDA
  • Experience building synthetic data pipelines for agent training
  • Familiarity with open-source RL frameworks: TRL, OpenRLHF, veRL/vLLM

Hourly Pay:  $35- $50 

The US base salary range for this full-time position is listed below. Pay may vary based on a number of factors including job-related skills, level, experience, geographic location and relevant education or training. At NewsBreak, we design our overall rewards package to attract top talents. Depending on the position, the role may also be eligible for discretionary bonus and options. Your recruiter can share more details during the hiring process.
Annual Base Pay Range
$35$50 USD
Ready to apply to NewsBreak?
Apply to NewsBreak

Similar jobs

NewsBreak
Local Intent Data Mining Specialist
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $112,000–$293,000
● New 👁 Seen ✓ Applied 1d ago
NewsBreak
Localization Vertical Product Manager
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $146,000–$342,000
● New 👁 Seen ✓ Applied 1d ago
NewsBreak
Senior Data Infra Engineer
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $175,000–$221,000
● New 👁 Seen ✓ Applied 4d ago
NewsBreak
Local POI Knowledge Graph
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $142,000–$342,000
● New 👁 Seen ✓ Applied 5d ago
NewsBreak
AI Engineer, Agentic Ad Creative (Multimodal)
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $120,000–$220,000
● New 👁 Seen ✓ Applied 6d ago
NewsBreak
Matching Algorithm Engineer
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $163,000–$400,000
● New 👁 Seen ✓ Applied 1w ago
NewsBreak
Senior Machine Learning Engineer, User Signal & Ads
NewsBreak
⚡ Apply early Bellevue, Washington, United S... Onsite $185,000–$235,000
● New 👁 Seen ✓ Applied 1w ago
NewsBreak
Applied AI Engineer, Advertising Agents
NewsBreak
⚡ Apply early Mountain View, California, Uni... Onsite $135,000–$185,000
● New 👁 Seen ✓ Applied 1w ago
NewsBreak
Growth Intelligence Engineer (Ads & Revenue)
NewsBreak
⚡ Apply early Mountain View, California, Uni... Hybrid $145,000–$185,000
● New 👁 Seen ✓ Applied 1w ago

Sign up for suggestions tailored to the jobs you open and the searches you save.

Apply now
🤖

Whoa — hold up

JobsRadar was built for real people having a rough time in their job search — not for automated requests. You're clicking way too fast and you're now temporarily blocked.

Come back later. If you're genuinely job hunting, we've got your back — just act like a human.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Get the worldwide-remote edge.

Join our Telegram channel for the stuff that helps you land the role — salary benchmarks, the weekly market pulse, and new-feature drops. No spam, just signal.

Join the channel — it's free