About this Data Scientist role at DevSavant Inc.
About the DevSavant
DevSavant is an operating partner for startups and growth-stage companies, helping them turn ambition into execution.
We support founders and leadership teams with product engineering and global staffing, from early prototypes and MVPs to scaling high-performing teams. Our vetted talent across LATAM and Asia embeds directly into client teams, operating as true extensions rather than external vendors.
With over 8 years working in venture-backed ecosystems, DevSavant is trusted to accelerate delivery, scale teams efficiently, and support companies as they reach their next milestone.
About the Role
We're looking for a Data Scientist to build the models and data products that make it all the way to production — from generative models on mixed data sources, to subscriber-behavior predictions, to new models for TV providers (MVPDs). You'll dig into large, messy datasets to find the trends and patterns that turn into shipped features, and you'll build LLM-powered pipelines and agents with the evals to prove they work. You'll work closely with data scientists and engineers to take ideas from first experiment to production at market scale, and collaborate directly with cross-functional stakeholders — including our co-founders.
This is a mid-level, remote role reporting into the Data Science team, requiring advanced (C1) English proficiency for clear, direct communication on complex technical and system design decisions.
Key Responsibilities
Build models and data products that make it all the way to production — from generative models on mixed data sources to subscriber-behavior predictions and new models for TV providers (MVPDs).
Dig into large, messy datasets to find the trends and patterns that turn into shipped features, and add the functions, classes, and tools to our core Python data science library that the rest of the team builds on.
Take on Antenna R&D work: explore new datasets and methods to answer real business questions and present what you find to senior stakeholders.
Write clear, well-organized, testable, and efficient code using object-oriented principles, grounded in deep knowledge of core Python and data tools. Because you care about quality, your code is well-documented.
Debug complex distributed systems and make code faster and able to handle more data.
Build LLM-powered pipelines and agents, and explain the failure modes you hit and the guardrails you added.
Treat evals as a core deliverable: validate model responses with provider-enforced structured outputs, build eval sets with clear pass/fail checks, calibrate LLM-as-a-judge rubrics, and use tracing tools to track cost, latency, and quality over time. You can point to an eval that caught a problem human review missed.
Use agentic coding tools as part of your daily workflow: plan first, write tests and instructions up front, and review every change before accepting it — while still designing, debugging, and defending your work without AI assistance.
Collaborate with cross-functional stakeholders, including our co-founders, and clearly explain complex technical and system design decisions.
Required Qualifications
2+ years of experience building machine learning models and data products in Python, with the engineering skills to take them from prototype to production.
Expert in Python with strong object-oriented design, software system design, and experience building high-quality, testable, production-grade code.
Hands-on experience with deep learning frameworks (PyTorch or TensorFlow), plus a deep understanding of machine learning concepts, the end-to-end model development lifecycle, and MLOps principles.
Hands-on experience with large-scale data processing tools (e.g., Apache Spark/PySpark, Dask) and strong SQL skills working with large, complex datasets.
Solid experience with cloud platforms (GCP highly preferred), including deploying, managing, and scaling services (Docker, Cloud Run, GKE) and working with big data systems (Dataproc, BigQuery).
Excellent problem-solver, skilled at debugging complex distributed systems and optimizing them for performance and scale.
Advanced English proficiency (B2–C1) with strong communication, teamwork, and consulting skills; able to clearly explain complex technical and system design decisions.
Daily use of agentic coding tools (Claude Code, Cursor, or Codex CLI): plan first, write tests and instructions up front, and review every change before accepting it. You can still design, debug, and defend your own work without AI assistance — our interview process tests this directly.
Experience building and shipping LLM-powered agents or pipelines using an orchestration framework (LangGraph, Pydantic AI, or OpenAI Agents SDK), including custom tool definitions against internal APIs and data systems, agent state and memory, and human review steps. You can explain the failure modes you hit and the guardrails you added.
You treat evals as a core deliverable: you validate model responses with provider-enforced structured outputs (Pydantic), build eval sets with clear pass/fail checks, calibrate LLM-as-a-judge rubrics, and use tracing tools (Langfuse, LangSmith, or Braintrust) to track cost, latency, and quality over time. You can describe an eval that caught a problem a human review missed.
Bonus
Experience in or passion for the Subscription Economy, especially media and entertainment; experience working with media data or data clean rooms is a plus.
Experience with synthetic data generation or advanced generative models (e.g., GANs, VAEs, CTGAN).
Experience building Python libraries that others use, or contributions to open-source projects.
Knowledge of advanced MLOps practices (like model monitoring) and build automation tools (e.g., Cloud Build, Cloud Run).
Experience building custom tool integrations for agents: your own tool definitions and routing against internal APIs and data systems, with clear input schemas, validation, and safe handling of side-effecting actions.
Experience with advanced evaluation and observability practices like multi-judge calibration, automated regression suites, and production monitoring of agent quality.
Familiarity with RAG and context engineering for grounding model responses in proprietary data.
Experience using LLMs for testing pipelines and QA workflows.