Jobs › Companies › Alpaca › Senior Data Scientist AI Evaluation

Über diese Senior Data Scientist AI Evaluation Stelle bei Alpaca

Alpaca · Remote · Remote - Americas

Who We Are:

Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.

Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.

Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.

Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.

Our Team Members:

We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!

We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.

Your Role: We're looking for a Senior Data Scientist, AI Evaluation to design how Alpaca measures whether our models and agents are actually right. You'll be a senior individual contributor who turns ambiguous quality questions into ground truth, scoring methods, and eval loops that the company can trust—and uses those results to make the systems better. You'll build on an established data foundation, so the focus is raising quality and speeding up safe rollout. You'll own the quality bar, independent of the teams that build and optimize those systems.

 

This role is for someone who cares as much about whether an answer is correct as about whether a model can generate one. You'll partner with Product, Engineering, Analytics Engineering, and business stakeholders to define what "good" looks like, build the evaluations that test it, and close the loop so evals drive iteration. If you have a strong quantitative background, have shipped rigorous, measurable work (evaluation, experimentation, or model validation), and want ownership over a greenfield eval practice at a fast-growing brokerage-infrastructure company, this is the role.

 

What You'll Do

  • Design AI evaluations: Define ground truth, metrics, and scoring methods for models and agents.
  • Build repeatable eval loops: Track quality over time and catch regressions before release.
  • Partner on infrastructure: Work with engineering and analytics engineering to operationalize eval harnesses.
  • Drive iteration: Translate eval results into actionable recommendations for system improvements.
  • Establish quality standards: Set evaluation guidelines, documentation, and review practices.
  • Mentor and align: Foster evaluation best practices and build a culture of measurable AI quality across the team.

What We're Looking For

  • Track record of quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation).
  • Strong statistical and ML foundation—you treat evaluations as experiments (sample sizing, confidence intervals, significance, handling non-determinism) and validate automated graders against human ground truth.
  • Proficiency in Python and SQL, with experience evaluating models in production environments.
  • Strong judgment in defining quality metrics and ground truth for ambiguous outputs.
  • Excellent communication and cross-functional collaboration skills to align technical teams and leadership.
  • Strong problem-solving ability in fast-paced, greenfield environments.
  • 6–10 years in quantitative data science or ML, with focused experience in measurement or evaluation. A quantitative degree is a plus; equivalent industry experience is equally welcome.

Nice to Have

  • Hands-on LLM/agent evaluation in production, including eval harnesses, LLM-as-judge calibration, and CI regression gates.
  • Experience evaluating text-to-SQL, analytics agents, or other systems where correctness is verifiable against data.
  • Background in fintech, brokerage, or other domains where a wrong answer has real business or risk consequences.
  • Fluency with AI tools in research and engineering workflows.

How We Take Care of You:

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

Recruitment Privacy Policy

Bereit, sich bei Alpaca zu bewerben?
Bei Alpaca bewerben

Über Alpaca

Who We Are:

Alpaca is a US California headquartered brokerage infrastructure technology company and self-clearing broker-dealer, delivering execution and custody solutions for Stocks, ETFs, Options, Cryptocurrencies, and more, and has raised over $170 million in funding. Amongst our subsidiaries, Alpaca is a licensed financial services company in multiple countries, and we serve hundreds of financial institutions globally such as broker-dealers, investment advisors, hedge funds, and crypto exchanges as well as millions of individual customers all over the world.

 

Alpaca’s globally distributed team members bring in diverse experiences such as engineers, traders, and brokerage professionals to achieve our Mission of opening financial services to everyone on the planet. We are also deeply committed to open-source contributions and fostering a vibrant community. We will continue to enhance and improve our award-winning developer-friendly API and the brokerage infrastructure behind it.



Our Team Members:

We’re a team of 200+ globally distributed members who love working from our favorite places worldwide. Our team spans the USA, Canada, Japan, Hungary, Nigeria, Brazil, the United Kingdom, and more!

We’re looking for candidates eager to join Alpaca’s growing organization, who are excited about our Mission of “Open[ing] financial services to everyone on the planet” and share our Values of “Stay Curious,” “Have Empathy,” and “Be Accountable.”

Alle Jobs bei Alpaca ansehen →

Ähnliche Jobs

Jobber
Staff Data Scientist, Applied ML
Jobber
⚡ Früh bewerben Remote Hybrid CA$145,900–CA$145,900
● Neu 👁 Gesehen ✓ Beworben vor 2 Std.
Life360
Lead Data Scientist
Life360
⚡ Früh bewerben Remote, USA; Remote, Canada · standortgebunden $175,000–$218,000
● Neu 👁 Gesehen ✓ Beworben vor 5 Std.
ZoomInfo Technologies LLC
Principal Data Scientist - Product Analytics
ZoomInfo Technologies LLC
⚡ Früh bewerben Remote · standortgebunden $136,500–$214,500
● Neu 👁 Gesehen ✓ Beworben vor 5 Std.
Nebius
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius
⚡ Früh bewerben Amsterdam, Netherlands; Berlin... Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Oasis Health Partners
Principal Data Scientist
Oasis Health Partners
⚡ Früh bewerben Remote · standortgebunden
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Reddit
Principal Data Scientist, Ads
Reddit
⚡ Früh bewerben Remote - United States · standortgebunden $268,000–$365,100
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Reddit
Staff Data Scientist - Ads Measurement, Signals, Privacy
Reddit
⚡ Früh bewerben Toronto, Canada Vor Ort
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Reddit
Staff Data Scientist - Ads Measurement, Signals, Privacy
Reddit
⚡ Früh bewerben Remote - United States · standortgebunden $217,000–$303,900
● Neu 👁 Gesehen ✓ Beworben vor 8 Std.
Muttdata
Data Scientist Semi Senior - Databricks – Causal Inference & Growth Marketing #5
Muttdata
⚡ Früh bewerben Weltweit
● Neu 👁 Gesehen ✓ Beworben vor 9 Std.

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Alpaca

Alle Jobs bei Alpaca ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos