Z
Posted Yesterday•Remote
Senior Data Engineer
SeniorRemoteSalary undisclosed
Required Skills
PythonAngularAWSDockerTerraformPostgreSQLSQLFastAPILangChainDatabricks
Job Description
Design, build, and operate the data infrastructure for a clinical trial design platform that extracts insights from clinical protocols, regulatory documents, and published articles to accelerate trial design decisions. This role owns the pipelines that ingest, parse, transform, and serve clinical data - from raw PDF extraction through structured storage in Aurora and GraphDB, to serving curated knowledge for LangGraph-based agentic AI workflows.
Requirements
- Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization
- Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships
- Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population
- Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores
- Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.