ParetoHealth logo
Posted 3h agoRemote

Lead Data Scientist

LeadRemoteSalary undisclosed
Required Skills
PythonAWSSQLPyTorch
Job Description

About ParetoHealth

ParetoHealth is redefining the way employers fund healthcare. As the largest and fastest-growing benefits captive in the United States, we help thousands of small and midsize employers take control of healthcare costs through a smarter, more sustainable model.

Our mission is simple: give small and midsize employers the scale and protection they need to eliminate volatility and lower healthcare costs.

By combining data-driven insights, innovative risk management, and the collective purchasing power of our community, we enable employers to reduce volatility, improve long-term outcomes, and reinvest savings into their businesses and their people.

Headquartered in Philadelphia, ParetoHealth is growing rapidly and transforming one of the country's largest industries. Our success is fueled by talented people who are united by our four core values: Fire in the Belly, For the Greater Good, See the Field, and Get It Done Right. These values shape how we innovate, collaborate, make decisions, and deliver exceptional results for our clients and one another.

If you're energized by solving complex challenges, thrive in a high-growth environment, and want to help reshape the future of healthcare, we'd love to meet you.

Please note that ParetoHealth does not provide employment visa sponsorship for this position. Candidates must be authorized to work in the United States without sponsorship both now or in the future.

Position Summary:

The Lead Data Scientist will independently lead applied AI and data science workstreams that translate complex healthcare and underwriting data into measurable improvements in risk selection, pricing accuracy, operational efficiency, and the underwriting experience. The role will personally lead work from problem framing, target definition, feature engineering, and model development through rigorous validation, production deployment, monitoring, and business-impact measurement. The Lead Data Scientist will operate with broad autonomy, manage workstream plans and risks, and bring major methodological, governance, or cross-platform decisions to the VP of AI and VP Analytics when a decision has broader business impact.

In addition to delivering models, the Lead Data Scientist will help build internal AI products and contribute improvements to shared practices for temporal validation, explainability, responsible AI and data use, model governance, and performance monitoring. Working with the VP of AI and VP of Analytics, the role will align assigned workstreams to the broader scientific roadmap and escalate decisions with cross-workstream impact. The role will partner with Business, AI, and Engineering to translate complex evidence into clear recommendations and measurable business outcomes.

Key Responsibilities:

  • Own end-to-end delivery and measurable outcomes for assigned predictive underwriting and pricing workstreams, from problem framing through deployment, monitoring, and continuous improvement, including:
    • Frame evidence-based recommendations and workstream trade-offs across model quality, risk, cost, scalability, aligning partners on delivery and measurable outcomes
    • Defining and documenting the Analytics-Ready Dataset, data-quality, and reusable feature requirements needed for assigned workstreams across claims, pharmacy, utilization, financial, underwriting, and external data
    • Applying rigorous point-in-time development and out-of-time validation approaches for claims maturity, seasonality, leakage, stability, and uncertainty
    • Developing, comparing, and challenging predictive models, distributions, and hybrid rule/model approaches based on evidence and operating constraints
    • Balance near-term delivery with disciplined exploration of emerging methods, including deep learning or GenAI solutions, that can materially improve accuracy, scalability, decision quality, or operating efficiency
    • Translating model needs into detailed feature requirements; partnering with business teams to identify additional signals and with Legal to secure approvals
    • Managing third-party model evaluations and ROI analyses when external expertise or independent validation is needed
    • Producing explainable, reproducible model outputs and following shared documentation, testing, monitoring, retraining, and rollback standards while recommending improvements based on workstream experience
  • Champion reuse, standardization, and componentization of data science assets so successful features, pipelines, evaluation patterns, and scoring capabilities can scale across Pareto Predict
  • Translate model outputs into decision support that improves underwriting accuracy, efficiency, adoption, and user experience; use measured business results to recommend whether to scale, iterate, or stop
  • Ensure assigned solutions operate within shared evaluation, monitoring, model-risk, responsible-AI, privacy, and regulatory expectations
  • Contribute to team capability through hands-on technical reviews, reusable components, mentoring, and pragmatic evaluation of new statistical, machine-learning, and AI methods

Key Characteristics:

  • An independent workstream owner who combines strong scientific judgment with hands-on delivery
  • A strategic problem-solver who connects modeling decisions to underwriting outcomes and measurable economic value
  • An influential collaborator who aligns Underwriting, Business, Product, and Engineering without relying on formal authority
  • An effective communicator who explains technical trade-offs, risks, and business impact in a way that enables confident decisions across technical and business audiences
  • A collaborative builder who contributes reusable components, documentation, and practical improvements to shared standards

Required Skills & Qualifications:

  • Bachelor's or master's degree in Statistics, Data Science, Computer Science, Mathematics, Engineering, or a related quantitative field; an advanced degree is a plus
  • 8+ years in data science, machine learning, statistics, actuarial analytics, including substantial work with healthcare, pharmacy, insurance risk, or sensitive longitudinal data and a proven record of independently owning models or analytical workstreams end to end
  • Advanced Python and SQL, with experience in scikit-learn, XGBoost, GBMs, or comparable frameworks. Experience using PySpark or comparable distributed-computing tools to work with structured and unstructured data at scale is preferred; PyTorch experience is a plus
  • Deep expertise in Supervised and Unsupervised Machine Learning methods, explainability, statistical distributions, rare-event and high-cost modeling, calibration, and optimization
  • Proven ownership of production models and the MLOps lifecycle on AWS or a comparable cloud platform, including Git-based version control, testing, deployment, monitoring, retraining, rollback, documentation, and responsible AI/model-governance practices.
  • Familiarity with Kedro or a similar pipeline framework is a plus. Experience with LLMs, prompt engineering, RAG, embeddings, vector DB, or agentic frameworks is helpful but not required.
  • Strong business acumen and judgment, with the ability to connect analytical outputs to measurable business outcomes, influence cross-functional decisions, and navigate trade-offs among model quality, risk, cost, scalability, and time to value.

Perks & Benefits:

  • Fully paid medical, dental, and vision benefits.
  • Flexible PTO
  • 401k company contribution
  • Tuition reimbursement
  • Professional development allowance
  • Transportation allowance and daily parking reimbursement 
  • Engaging hybrid work environment

We are guided by our values:

Fire in the belly

The drive to learn, to improve, and to deliver outstanding value every day.

See the field

The ability to see the big picture and prepare to meet tomorrow’s needs.

Get it done right

The passion to produce at higher rates and to the highest standards.

For the greater good

A united community creating better health benefit solutions for all.

Please note that any communication from our recruiters and hiring managers at ParetoHealth about a job opportunity will only be made by a ParetoHealth employee with an @paretohealth.com address. ParetoHealth does not conduct text message or chat-based interviews. Any other email addresses, agencies, or forums may be phishing scams designed to obtain your personal information.

We will not ask you to provide personal or financial information, including, but not limited to, your social security number, online account passwords, credit card numbers, passport information, and other related banking information until we begin onboarding activities, which will be coordinated by a member of the ParetoHealth People Ops Team with an @paretohealth.com email address.

Disclosures:
ParetoHealth is an Equal Opportunity Employer and does not discriminate on the basis of race, color, religion (creed), gender, gender expression, age, national origin (ancestry), disability, marital status, sexual orientation, or military status, in any of its activities or operations. These activities include, but are not limited to, hiring and firing of staff, selection of volunteers and vendors, and provision of services. We are committed to providing an inclusive and welcoming environment for all members of our staff, clients, volunteers, subcontractors, vendors, and clients.
California Applicants:  See Pareto’s CCPA Notice of Collection for California Employees and Applicants for information about how Pareto Captive Services, LLC, Pareto Health, LLC, and Pareto Underwriting Partners, LLC, together with their respective subsidiaries (collectively, “Pareto”) collects and uses personal information submitted by employment applicants.

Similar Openings in AI & Machine Learning

View all in category