Sobre esta vaga de Lead Data Scientist na ParetoHealth
About ParetoHealth
ParetoHealth is redefining the way employers fund healthcare. As the largest and fastest-growing benefits captive in the United States, we help thousands of small and midsize employers take control of healthcare costs through a smarter, more sustainable model.
Our mission is simple: give small and midsize employers the scale and protection they need to eliminate volatility and lower healthcare costs.
By combining data-driven insights, innovative risk management, and the collective purchasing power of our community, we enable employers to reduce volatility, improve long-term outcomes, and reinvest savings into their businesses and their people.
Headquartered in Philadelphia, ParetoHealth is growing rapidly and transforming one of the country's largest industries. Our success is fueled by talented people who are united by our four core values: Fire in the Belly, For the Greater Good, See the Field, and Get It Done Right. These values shape how we innovate, collaborate, make decisions, and deliver exceptional results for our clients and one another.
If you're energized by solving complex challenges, thrive in a high-growth environment, and want to help reshape the future of healthcare, we'd love to meet you.
Please note that ParetoHealth does not provide employment visa sponsorship for this position. Candidates must be authorized to work in the United States without sponsorship both now or in the future.
Position Summary:
The Lead Data Scientist will independently lead applied AI and data science workstreams that translate complex healthcare and underwriting data into measurable improvements in risk selection, pricing accuracy, operational efficiency, and the underwriting experience. The role will personally lead work from problem framing, target definition, feature engineering, and model development through rigorous validation, production deployment, monitoring, and business-impact measurement. The Lead Data Scientist will operate with broad autonomy, manage workstream plans and risks, and bring major methodological, governance, or cross-platform decisions to the VP of AI and VP Analytics when a decision has broader business impact.
In addition to delivering models, the Lead Data Scientist will help build internal AI products and contribute improvements to shared practices for temporal validation, explainability, responsible AI and data use, model governance, and performance monitoring. Working with the VP of AI and VP of Analytics, the role will align assigned workstreams to the broader scientific roadmap and escalate decisions with cross-workstream impact. The role will partner with Business, AI, and Engineering to translate complex evidence into clear recommendations and measurable business outcomes.
Key Responsibilities:
- Own end-to-end delivery and measurable outcomes for assigned predictive underwriting and pricing workstreams, from problem framing through deployment, monitoring, and continuous improvement, including:
- Frame evidence-based recommendations and workstream trade-offs across model quality, risk, cost, scalability, aligning partners on delivery and measurable outcomes
- Defining and documenting the Analytics-Ready Dataset, data-quality, and reusable feature requirements needed for assigned workstreams across claims, pharmacy, utilization, financial, underwriting, and external data
- Applying rigorous point-in-time development and out-of-time validation approaches for claims maturity, seasonality, leakage, stability, and uncertainty
- Developing, comparing, and challenging predictive models, distributions, and hybrid rule/model approaches based on evidence and operating constraints
- Balance near-term delivery with disciplined exploration of emerging methods, including deep learning or GenAI solutions, that can materially improve accuracy, scalability, decision quality, or operating efficiency
- Translating model needs into detailed feature requirements; partnering with business teams to identify additional signals and with Legal to secure approvals
- Managing third-party model evaluations and ROI analyses when external expertise or independent validation is needed
- Producing explainable, reproducible model outputs and following shared documentation, testing, monitoring, retraining, and rollback standards while recommending improvements based on workstream experience
- Champion reuse, standardization, and componentization of data science assets so successful features, pipelines, evaluation patterns, and scoring capabilities can scale across Pareto Predict
- Translate model outputs into decision support that improves underwriting accuracy, efficiency, adoption, and user experience; use measured business results to recommend whether to scale, iterate, or stop
- Ensure assigned solutions operate within shared evaluation, monitoring, model-risk, responsible-AI, privacy, and regulatory expectations
- Contribute to team capability through hands-on technical reviews, reusable components, mentoring, and pragmatic evaluation of new statistical, machine-learning, and AI methods
Key Characteristics:
- An independent workstream owner who combines strong scientific judgment with hands-on delivery
- A strategic problem-solver who connects modeling decisions to underwriting outcomes and measurable economic value
- An influential collaborator who aligns Underwriting, Business, Product, and Engineering without relying on formal authority
- An effective communicator who explains technical trade-offs, risks, and business impact in a way that enables confident decisions across technical and business audiences
- A collaborative builder who contributes reusable components, documentation, and practical improvements to shared standards
Required Skills & Qualifications:
- Bachelor's or master's degree in Statistics, Data Science, Computer Science, Mathematics, Engineering, or a related quantitative field; an advanced degree is a plus
- 8+ years in data science, machine learning, statistics, actuarial analytics, including substantial work with healthcare, pharmacy, insurance risk, or sensitive longitudinal data and a proven record of independently owning models or analytical workstreams end to end
- Advanced Python and SQL, with experience in scikit-learn, XGBoost, GBMs, or comparable frameworks. Experience using PySpark or comparable distributed-computing tools to work with structured and unstructured data at scale is preferred; PyTorch experience is a plus
- Deep expertise in Supervised and Unsupervised Machine Learning methods, explainability, statistical distributions, rare-event and high-cost modeling, calibration, and optimization
- Proven ownership of production models and the MLOps lifecycle on AWS or a comparable cloud platform, including Git-based version control, testing, deployment, monitoring, retraining, rollback, documentation, and responsible AI/model-governance practices.
- Familiarity with Kedro or a similar pipeline framework is a plus. Experience with LLMs, prompt engineering, RAG, embeddings, vector DB, or agentic frameworks is helpful but not required.
- Strong business acumen and judgment, with the ability to connect analytical outputs to measurable business outcomes, influence cross-functional decisions, and navigate trade-offs among model quality, risk, cost, scalability, and time to value.
Perks & Benefits:
- Fully paid medical, dental, and vision benefits.
- Flexible PTO
- 401k company contribution
- Tuition reimbursement
- Professional development allowance
- Transportation allowance and daily parking reimbursement
- Engaging hybrid work environment
We are guided by our values:
Fire in the belly
The drive to learn, to improve, and to deliver outstanding value every day.
See the field
The ability to see the big picture and prepare to meet tomorrow’s needs.
Get it done right
The passion to produce at higher rates and to the highest standards.
For the greater good
A united community creating better health benefit solutions for all.