Über diese Data Scientist Stelle bei AffirmedRx, PBC
AffirmedRx is on a mission to improve health care outcomes by bringing clarity, integrity, and trust to pharmacy benefit management. We are committed to making pharmacy benefits easy to understand, straightforward to access and always in the best interest of employers and the lives they impact. We accomplish this by bringing total clarity to business practices, leading with clinical approaches, and utilizing state-of-the-art technology.
Join us in improving health care outcomes for all! We promise to do what’s right, always.
Position Summary:
The Data Scientist (AI/ML) designs, builds, and validates the advanced analytics that turn pharmacy, claims, clinical, and member data into decisions. This is a hands-on modeling role: the person owns machine learning models end to end, applies AI and natural language processing to unstructured clinical and member data, resolves member identity across fragmented data sources, and packages results into tools and dashboards the business can use. The role sits at the intersection of data science, clinical/pharmacy reporting, and applied AI, and partners closely with data engineering, clinical, reporting, and client-success teams.
What you will do:
Machine Learning and Predictive Modeling:
- Build predictive and prescriptive ML models for pharmacy cost and risk (e.g., forecasting second-year member spend), including feature engineering, model selection, and explainability analysis (e.g., SHAP-based feature attribution)
- Develop member-level risk and comorbidity scoring, mapping drug identifiers (NDC → ATC) to clinical conditions and severity weights, and validating outputs against edge cases
- Apply ML to automate high-effort clinical operations processes (e.g., prior-authorization override automation), moving manual workflows into rules-based and model-driven pipelines
Applied AI and Natural Language Processing:
- Use AI/NLP to analyze unstructured member and clinical text — sentiment analysis, topic modeling, and tokenization of open-ended survey and feedback data
- Apply AI tooling (LLMs / copilots and internal AI services) to automate clinical policy and documentation workflows, including prompt design, output validation, and controls against hallucination and format drift
- Contribute to the organization’s broader AI direction: evaluating models, defining evaluation/answer-key datasets, and building drift and validation checks for AI outputs
Member Identity Resolution and Data Quality:
- Design and maintain probabilistic (fuzzy) matching logic to assign and reconcile unique member identifiers across carriers and source systems, including collision handling, cluster analysis, and audit/logging frameworks
- Monitor and improve match rates, investigate false positives and fragmentation, and document data lineage and safeguards against duplicates
Clinical and Pharmacy Analytics:
- Produce clinical and pharmacy analytics such as medication adherence and persistence (drug-, class-, and NDC-level), aligned to compliance requirements (e.g., URAC / PQA measures)
- QA and validate reporting products (e.g., pharmacy trend dashboards, PMPM metrics), reconciling data-point discrepancies across source systems
Analytical Tooling and Delivery:
- Build analytical tools and prototypes (e.g., formulary/tier decision tools and cost-comparison tools), including lightweight front ends (e.g., Streamlit) for sales, clinical, and pricing use
- Deliver validated datasets and tables into the data warehouse in partnership with data engineering, and support the transition of prototypes into production
Validation, Documentation, and Collaboration:
- Own QA and validation for analytical outputs, including auditing of claims files and validation of model results before release
- Document models, logic, data sources, schedules, and troubleshooting steps to make work reproducible and auditable
- Collaborate across clinical, reporting, pricing, client-success, and engineering stakeholders to gather requirements and translate them into analytical specifications
What you need:
- Degree in a quantitative field (data science, statistics, computer science, applied math) or equivalent experience
- 2–3 years of experience using Python and SQL for data analysis, machine learning, NLP, data quality, and record-matching solutions. Data modeling experience in PBM and/or healthcare industry in general preferred
- Strong Python for data science and ML (e.g., pandas plus a modeling stack), and proficiency in SQL
- Demonstrated experience building and validating ML models, including feature engineering and model explainability
- Experience with NLP techniques (sentiment analysis, topic modeling) and with applying AI/LLM tooling to real workflows, including output validation
- Experience with entity resolution / probabilistic record matching and data-quality analysis
- Comfort working with a modern cloud data warehouse and data lake, and partnering with data engineering on production hand-off
- Healthcare, pharmacy benefit management (PBM), or claims-data experience
- Familiarity with pharmacy data concepts (NDC, GPI, ATC, formulary tiers, prior authorization, rebates)
- Experience with compliance-driven reporting (e.g., URAC / PQA measures)
- Experience building analytical front ends or dashboards (e.g., Streamlit, BI tools) for non-technical stakeholders
- Willingness and ability to travel (10%-20%)
What you get:
- To impact industry change in the pharmacy benefits management space, while delivering the highest quality patient outcomes
- To work in a culture where people thrive because when OUR team thrives, OUR business thrives
- Competitive compensation, including health, dental, vision and other benefits
Note:
AffirmedRx is committed to providing equal employment opportunities to all employees and applicants for employment. Remote employees are expected to maintain a professional work environment free of distractions to ensure optimal performance and collaboration.