Sobre este puesto de AI Engineer Internship – LLM Data en Ifm Us
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The Role
As an AI Engineer Intern specializing in LLM data, you will work with our team to build and improve high-quality training data for foundation models across pre-training, mid-training, and post-training. You will gain hands-on experience with large-scale data processing, LLM-based data synthesis, data quality evaluation, and model experimentation.
Responsibilities
-
Build, curate, and improve datasets for LLM pre-training, mid-training, and post-training, including time-sensitive requests that may require fast turnaround.
-
Develop and improve data processing workflows for data extraction, cleaning, filtering, deduplication, transformation, and quality control.
-
Use LLMs to generate, refine, filter, and evaluate synthetic data.
-
Analyze data quality and identify issues such as duplication, low-quality samples, distribution gaps, and coverage limitations.
-
Run experiments to understand how different data sources and processing strategies impact model performance.
-
Support model inference, evaluation, and training experiments when needed to validate data quality.
-
Research new datasets, data processing techniques, and LLM data methodologies.
-
Collaborate closely with researchers and engineers on fast-moving foundation model projects.
Qualifications
-
Currently pursuing a Bachelor’s, Master’s, or PhD degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related technical field.
-
Strong programming skills in Python.
-
Good understanding of machine learning, deep learning, NLP, or large language models.
-
Hands-on experience with LLMs through research, coursework, projects, internships, or open-source work.
-
Comfortable working with datasets and performing data processing and analysis.
-
Strong problem-solving skills and willingness to learn new technologies quickly.
-
Ability to work independently and collaborate effectively in a fast-paced research environment.
Preferred Qualifications
-
Experience with PyTorch, Hugging Face, vLLM, or similar ML/LLM frameworks.
-
Experience with LLM data synthesis, fine-tuning, model evaluation, or prompt-based generation.
-
Familiarity with LLM training pipelines, including pre-training, supervised fine-tuning, or other post-training methods.
-
Experience working with large-scale datasets or distributed data processing.
-
Relevant research, open-source contributions, competitions, or personal projects in LLMs or generative AI.