Sobre esta vaga de Senior Research Engineer na Weekday AI
This role is for one of Weekday’s clients
Salary range: Rs 5000000 - Rs 9000000 (ie INR 50 - 90 LPA)
Min Experience: 1+ years
Location: Bengaluru
JobType: full-time
We are looking for a highly motivated Senior Research Engineer with 1–8 years of experience to join our research and engineering team. The ideal candidate will work at the intersection of LLM research, model evaluation, and large-scale experimentation, helping develop rigorous methods to understand, measure, and improve the capabilities of modern language models.
You will design and implement evaluation frameworks, build high-quality datasets and test suites, analyze model behavior, and translate research findings into practical improvements. This role is ideal for someone who enjoys solving open-ended research problems while also being comfortable building production-quality systems.
Requirements
Key Responsibilities
- Design, develop, and maintain comprehensive LLM evaluation (evals) frameworks to measure model capabilities, reliability, reasoning, instruction following, safety, and task performance.
- Conduct research on large language models, including model behavior, capabilities, limitations, prompting, fine-tuning, and evaluation methodologies.
- Develop novel evaluation methodologies and experiments for emerging LLM capabilities and use cases.
- Create high-quality evaluation datasets, test cases, rubrics, and automated evaluation pipelines.
- Analyze model outputs using quantitative and qualitative methods to identify performance gaps and behavioral patterns.
- Design controlled experiments to compare models, prompts, training approaches, and inference strategies.
- Build scalable tooling for running evaluations across large numbers of prompts, models, and datasets.
- Collaborate with researchers, ML engineers, and product teams to convert research insights into measurable model improvements.
- Investigate failures and edge cases and develop targeted evaluations to capture previously undetected model weaknesses.
- Contribute to technical documentation, research reports, internal benchmarks, and presentations of findings.
- Stay current with developments in LLM research, evaluation techniques, reasoning systems, and AI benchmarks.
Required Skills & Qualifications
- 1–8 years of experience in machine learning, AI research, software engineering, data science, or a related technical field.
- Strong hands-on experience designing and implementing LLM evals or model evaluation systems.
- Solid understanding of LLM research, including model capabilities, prompting, fine-tuning, inference, and evaluation methodologies.
- Strong Python programming and experience working with ML/AI frameworks and data-processing pipelines.
- Ability to formulate research questions, design experiments, interpret results, and communicate technical findings clearly.
- Strong analytical and problem-solving skills with attention to experimental rigor and reproducibility.
- Experience working with large datasets, automated testing, and evaluation pipelines.
Good-to-Have Skills
- Experience developing or working with LLM benchmarks and standardized evaluation suites.
- Familiarity with benchmark design, dataset curation, scoring methodologies, and statistical analysis.
- Experience with open-source LLMs, model APIs, Hugging Face, PyTorch, or similar frameworks.
- Exposure to reinforcement learning, RLHF/RLAIF, fine-tuning, synthetic data generation, or agentic systems.
- Research publications, technical blogs, open-source contributions, or demonstrated independent AI research work.