Über diese Data Scientist Stelle bei Sambatv
As a Junior Data Scientist at Samba in Amsterdam, you will own end-to-end delivery of significant data science projects with minimal guidance. You are a reliable, autonomous contributor with deep expertise in at least one of Samba's core domains — Identity and audience modelling and the technical range to build production-ready solutions using modern ML and AI methodologies. You'll work closely with peers, product, and engineering.
What You'll Do:
-
Own end-to-end delivery of significant data science projects from problem scoping and approach design through to production deployment
-
Make sound, independently-reasoned decisions on methodology, model selection, and evaluation; document them clearly in technical solution documents covering problem statement, approach, metrics, and timeline
-
Lead solution design for your own initiatives; break down complex epics into well-scoped user stories with clear acceptance criteria, adopting DataOps and MLOps best practices throughout — experiment tracking, pipeline orchestration, model monitoring, and reproducibility
-
Implement advanced ML and AI-powered workflows including entity resolution, probabilistic record linkage, embedding-based matching, semantic similarity, and LLM-augmented pipelines
-
Develop and maintain reusable tools, libraries, and documentation that improve team efficiency and technical standards; conduct code reviews with constructive, specific feedback that raises the bar
-
Be a team player and collaborate with different teams.
-
Collaborate cross-functionally with product, engineering, and operations — translate business requirements into technical specifications, partner with data engineering on scalable pipeline design, and participate in cross-functional design reviews and working groups
Who You Are:
-
Bachelor's degree required in Statistics, Data Science, Computer Science, Mathematics or a related quantitative field; Master's strongly preferred
-
2+ years of hands-on data science experience with demonstrated ability to own and deliver complex, multi-sprint projects independently
-
Advanced Python (or any other programming language+Claude/Codex/any other) with production-quality code, testing, and documentation; strong SQL and PySpark for billion-row datasets
-
Databricks (nice to have/learn for 2027) workflows, Delta Lake, and job orchestration; working knowledge of cloud platforms (AWS or GCP)
-
Solid command of core ML — regression, classification, clustering, model evaluation, and experimental design — applied to complex, high-volume data (2 questions to level)
-
Proficiency with MLOps practices: experiment tracking and reproducible model deployment
-
Exposure to modern AI methodologies: RAG systems, LLM-augmented models, vector databases, and semantic search (showcase+examples)
-
Strong communicator — able to translate technical work into clear documentation, user stories, and cross-functional conversations
-
Demonstrated ability to mentor junior data scientists and contribute to team standards
-
Team player