Introduction
The Lead Data Scientist will lead data-driven initiatives that generate actionable insights, develop predictive and statistical models, and enable data-centric decision-making across healthcare business domains.
This role will help create well-governed master data products that stakeholders and end users can trust.
As the primary subject-matter contact for entity resolution, the Lead Data Scientist will address complex data-matching challenges across large, disparate data silos and collaborate with cross-functional teams to unlock insights from healthcare datasets.
Responsibilities
- Analyze large, complex datasets to identify trends, patterns, and actionable insights.
- Develop, implement, and validate statistical and machine learning models.
- Solve enterprise-scale record linkage, deterministic and probabilistic matching, and graph-based entity resolution challenges.
- Conduct root-cause analysis on training and test data to identify and resolve model anomalies.
- Establish robust validation and testing procedures to ensure data quality, model reliability, and data integrity.
- Support model training, inference, deployment, and ongoing performance monitoring.
- Develop and refine key performance indicators to measure success across business functions.
- Establish processes for tracking, reporting, and communicating performance metrics.
- Design and build data pipelines that support semantic-layer development and stakeholder needs.
- Contribute to trusted, scalable, and well-governed master data products.
Requirements
Required Skills: Python, SQL, .
NET, AWS, Databricks, Github, Healthcare Data
Preferred Skills: Machine Learning (ML)