Z
Posted Yesterday•Remote

Senior Data Engineer

SeniorRemoteSalary undisclosed
Required Skills
PythonAngularAWSDockerTerraformPostgreSQLSQLFastAPILangChainDatabricks
Job Description

Design, build, and operate the data infrastructure for a clinical trial design platform that extracts insights from clinical protocols, regulatory documents, and published articles to accelerate trial design decisions. This role owns the pipelines that ingest, parse, transform, and serve clinical data - from raw PDF extraction through structured storage in Aurora and GraphDB, to serving curated knowledge for LangGraph-based agentic AI workflows.

Requirements

  • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization
  • Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships
  • Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population
  • Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores
  • Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity

Similar Openings in Data & Analytics

View all in category➔