Jobs Companies MeridianLink Data Scientist, AI Data Foundations

À propos de ce poste Data Scientist, AI Data Foundations chez MeridianLink

MeridianLink · US Remote

Data Scientist, AI Data Foundations

About the Role

Reporting into the Data Engineering organization, the Data Scientist is responsible for designing and building the curated data structures that AI and ML applications consume across MeridianLink. You will own the vector stores behind our RAG systems, the feature store that powers model training and inference, and the graph databases that capture relationships across applicants, products, and decisions. You will also lead targeted data discovery work, surfacing hidden trends in our lending and account-opening data that inform both AI use cases and the broader business.

This is a hands-on, build-oriented role. You will not be primarily training large models — you will make sure the people training and serving models have high-quality, well-governed, well-engineered data to work with, and you will use your data science skills to validate that the data is fit for purpose.

What You Will Do

• Build and maintain vector stores for RAG: Design embedding pipelines, chunking strategies, indexing approaches, and refresh patterns for the vector stores powering retrieval-augmented generation across MeridianLink products.

• Own the feature store: Design, build, and operate feature store assets used for model training and online/offline inference, including feature definitions, freshness SLAs, lineage, point-in-time correctness, and reuse across teams.

• Design graph data structures: Build graph databases that model relationships between applicants, applications, products, lenders, decisions, and outcomes — and make them queryable for both AI use cases and analytical investigations.

• Lead data discovery: Profile our lending, deposit, and behavioral datasets to identify hidden trends, segments, anomalies, and potential model drivers; turn findings into actionable hypotheses for product, risk, and growth teams.

• Engineer for AI consumption: Build the curated, AI-ready datasets that downstream model builders, application engineers, and analysts rely on — with appropriate quality, documentation, and governance baked in.

• Evaluate retrieval and feature quality: Define and run evaluation frameworks for RAG retrieval quality, feature drift, embedding quality, and graph completeness; iterate based on what the metrics tell you.

• Partner with model builders: Work closely with ML engineers and applied scientists to make sure the data structures you build accelerate their work rather than slow it down.

• Champion responsible data use: Partner with governance, security, and compliance to ensure that AI-facing data assets respect data classification, customer consent, and regulatory boundaries from day one.

• Communicate findings: Translate discovery work into clear narratives — write-ups, notebooks, dashboards, and short presentations — that help non-technical stakeholders act on what the data is showing.

Required Qualifications

• 4–7 years of experience in a data science, ML engineering, or applied data role, with a meaningful portion of that time spent building data assets that other people's models or applications consumed.

• Hands-on experience designing and operating vector stores for RAG or semantic search, including embedding generation, chunking, indexing, and retrieval evaluation.

• Experience building or operating a feature store (e.g., Databricks Feature Store, Feast, or a custom internal platform), including offline training and online serving patterns and point-in-time correctness.

• Experience modeling and building graph data structures using Neo4j, TigerGraph, Azure Cosmos DB Gremlin, or similar graph databases — and writing graph queries to answer real questions.

• Strong proficiency in Python (pandas, NumPy, scikit-learn, PySpark) and SQL; comfortable working day-to-day in Databricks notebooks and jobs.

• Practical experience with embedding models and LLM tooling (e.g., Hugging Face transformers, OpenAI / Azure OpenAI APIs, LangChain or similar) in a production or near-production context.

• Demonstrated data discovery skills: profiling messy real-world datasets, surfacing non-obvious patterns, validating findings statistically, and explaining them clearly.

• Solid grounding in classical ML concepts — supervised vs. unsupervised learning, train/test discipline, leakage, evaluation metrics — even though you will not own model training day-to-day.

• Strong written and verbal communication skills; able to write up findings for both technical and business audiences.

Preferred Qualifications

• Experience working in a SaaS or FinTech environment, particularly with lending, deposit, credit, fraud, or KYC/AML data.

• Experience with Databricks-native AI/ML tooling: Databricks Vector Search, Databricks Feature Store, MLflow, and Unity Catalog.

• Familiarity with open-source vector databases such as pgvector, Pinecone, Weaviate, Chroma, or FAISS, and a clear point of view on when to use which.

• Experience with Microsoft Azure data and AI services (Azure OpenAI, Azure AI Search, ADLS Gen2).

• Experience evaluating RAG systems end-to-end (recall@k, faithfulness, answer quality, hallucination measurement).

• Exposure to graph algorithms (community detection, link prediction, centrality) applied to real business problems.

• Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field, or equivalent professional experience.

Our Data & AI Stack

• Lakehouse: Azure Databricks, Delta Lake, Unity Catalog, PySpark, SQL

• AI Data Foundations: Databricks Vector Search, Databricks Feature Store, MLflow

• Vector & Graph (current and exploratory): pgvector, Pinecone, Weaviate, FAISS; Neo4j, TigerGraph, Azure Cosmos DB (Gremlin)

• Cloud: Microsoft Azure (ADLS Gen2, Azure OpenAI, Azure AI Search, Event Hubs)

• AI Models and Agents: Databricks, AWS Bedrock, Azure ML

• Integration & Governance: Informatica Data Management Cloud (IDMC), Unity Catalog

Prêt à postuler chez MeridianLink ?
Postuler chez MeridianLink

Comment se compare ce salaire pour Data Scientist

Ce poste paie $144,750/yren dessous de la fourchette habituelle pour les postes Data Scientist.

$142,375 la médiane $191,600 $251,887

Fourchette typique $155,504–$213,757/yr, à partir de 47 annonces Data Scientist comparables sur JobsRadar (rémunération annualisée en USD). Voir les aperçus de salaire pour Data Scientist →

Emplois similaires

Block
Senior Data Scientist, AI & Model Risk
Block
⚡ Postuler tôt New York, NY, United States of... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 23 h
Block
Data Scientist
Block
⚡ Postuler tôt Bay Area, CA, United States of... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 23 h
Block
Senior Data Scientist, AI & Model Risk - Remote, US
Block
⚡ Postuler tôt Bay Area, CA, United States of... Sur site
● Nouveau 👁 Vu ✓ Postulé il y a 23 h
Mozilla
Senior Staff Data Scientist
Mozilla
⚡ Postuler tôt Remote; Remote Canada; Remote... · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Mozilla
Staff Data Scientist, Firefox
Mozilla
⚡ Postuler tôt Remote; Remote Canada; Remote... · lieu restreint
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Mozilla
Staff Data Scientist, Firefox
Mozilla
⚡ Postuler tôt Remote Canada; Remote US · lieu restreint $163,000–$218,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
Shift Technology
Data Scientist / Engineer (Healthcare Payment Integrity)
Shift Technology
⚡ Postuler tôt US - Remote · lieu restreint $120,000–$130,000
● Nouveau 👁 Vu ✓ Postulé il y a 1 j
EX
Senior Data Scientist, Revenue Analytics
Extend
⚡ Postuler tôt Remote, US · lieu restreint $125,000–$135,000
● Nouveau 👁 Vu ✓ Postulé il y a 4 j
Pinterest
Staff Data Scientist, Forecasting
Pinterest
⚡ Postuler tôt San Francisco, CA, US; Remote,... · lieu restreint $164,695–$339,078
● Nouveau 👁 Vu ✓ Postulé il y a 4 j

Inscrivez-vous pour des suggestions adaptées aux emplois que vous ouvrez et aux recherches que vous enregistrez.

Plus d’emplois chez MeridianLink

Voir tous les emplois chez MeridianLink →

Postuler maintenant
🤖

Doucement — un instant

JobsRadar a été conçu pour de vraies personnes qui traversent une période difficile dans leur recherche d’emploi — pas pour des requêtes automatisées. Vous cliquez beaucoup trop vite et vous êtes maintenant temporairement bloqué.

Revenez plus tard. Si vous cherchez réellement un emploi, nous sommes de votre côté — agissez simplement comme un être humain.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Prenez une longueur d’avance dans votre recherche d’emploi.

Rejoignez notre canal Telegram pour ce qui vous aide à décrocher le poste — références salariales, le pouls hebdomadaire du marché et les annonces de nouveautés. Pas de spam, que du signal.

Rejoindre le canal — c’est gratuit