G
Posted 2w agoPoland | Czechia | Croatia | Romania | Bulgaria | Hungary | Remote

Senior Data & Document Ingestion Engineer (OCR / RAG) - REMOTE

SeniorRemote~€1,600 – €2,200 / mo
Required Skills
PythonCI/CDData Engineering
Job Description

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About the Role

Our client is a Big4 Consultancy group that works with leading financial institutions on AI-driven transformation, automation, advanced analytics, and financial crime prevention. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.

We are looking for a Senior Data & Document Ingestion Engineer to build robust pipelines for processing high-volume unstructured insurance content. The role focuses on OCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation for downstream AI systems.

CONTRACT: Contractor assignment, expected October 2026 – July 2027, with extension to other projects (and retention rate)

COMMITMENT: Full-time

LOCATIONS: REMOTE 100%, Europe-based

PROCESS: Initial qualification followed by technical and client interviews

NOTES: Fluent English is required. Must be able to work in EU.

Responsibilities

  • Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents.
  • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction.
  • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
  • Process unstructured formats including PDF, Word, Excel, and PowerPoint.
  • Develop connectors for enterprise sources such as SharePoint and email systems.
  • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases.
  • Build validation and monitoring loops to detect low-confidence OCR or extraction results.
  • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
  • Implement logging, testing, and operational monitoring across data-processing workflows.
  • Apply Git, CI/CD, and software-engineering best practices to pipeline development.
  • Approximately 5–10 years of professional data engineering or backend/data-platform experience.
  • Strong hands-on experience with Python and SQL.
  • Proven experience building data ingestion and document-processing pipelines.
  • Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails.
  • Experience with OCR/document extraction tools such as AWS Textract or equivalent.
  • Professional experience building data-processing pipelines on public cloud platforms.
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services.
  • Strong development practices including Git, CI/CD, and automated testing.

Preferred Qualifications

  • Experience with Azure, AWS, or Databricks in enterprise data environments.
  • Experience with vector databases, embeddings, or RAG architectures.
  • Experience designing connectors to SharePoint, email, or other enterprise content systems.
  • Background in insurance, financial services, or regulated-data environments.
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.

Similar Openings in Backend

View all in category