Senior Data Engineer (AWS / Databricks)
Remote (EU/Ukraine) | Full-time
We’re hiring on behalf of our client — an international product company building and scaling a portfolio of subscription-based digital products for global markets.
The company is now building a centralized data platform that will bring together fragmented product, payments, marketing, and operational data across its portfolio. They are looking for a hands-on Senior Data Engineer to help establish Databricks on AWS, define the platform’s core engineering standards, and build reliable data products for analysts and business stakeholders.
This is a greenfield platform role with substantial technical ownership. You will not be joining a mature data environment with established patterns. You will help design those patterns, make foundational technical decisions, and create a repeatable approach for onboarding new products and data sources.
Why this role is interesting
Build a centralized data platform from an early stage rather than inherit a mature warehouse
Influence architecture, engineering standards, ingestion patterns, and governance
Solve a complex platform challenge involving distributed PostgreSQL databases across private AWS and EKS environments
Build the first portfolio-wide data models around payments, subscriptions, revenue, churn, LTV, and CAC
Work closely with Data, DevOps, Product, Backend Engineering, and business stakeholders
See a direct connection between your engineering work and key product and commercial decisions
What you’ll do
Build the Data Platform
Build and operate a Databricks-based data platform on AWS together with the Data and DevOps teams
Design and maintain Bronze, Silver, and Gold data layers using S3 and Delta Lake
Develop reusable ingestion patterns for PostgreSQL databases, S3, APIs, webhooks, and SaaS platforms
Build and manage production workflows using Databricks Jobs and Workflows
Contribute infrastructure changes through Terraform, Git, and pull-request-based workflows
Help establish platform standards, development patterns, and technical documentation
Build Reliable Data Pipelines
Implement incremental data loads, historical backfills, idempotent reprocessing, and schema-change handling
Design safe ingestion from multiple production PostgreSQL databases without creating unnecessary risk or load for source applications
Handle late-arriving updates, deletes, retries, and pipeline recovery
Build monitoring, freshness checks, reconciliation processes, and data-quality controls
Troubleshoot pipeline failures and data inconsistencies across multiple products and source systems
Optimize Databricks compute, SQL workloads, and storage for performance, reliability, and cost
Unify Product and Payments Data
Standardize fragmented product and payments data across the company’s portfolio
Build common analytical entities for users, subscriptions, transactions, renewals, refunds, and chargebacks
Normalize product-specific schemas into reliable source-of-truth models
Deliver trusted Gold datasets and data marts for Payments, Marketing, Product, Finance, and executive reporting
Support analytical use cases related to revenue, subscriptions, churn, LTV, CAC, product funnels, and attribution
Establish Governance and Engineering Standards
Contribute to Unity Catalog implementation and ongoing governance
Help manage groups, permissions, service principals, and data access patterns
Apply Git-based development, code review, CI/CD, testing, and documentation practices
Work with Product and Backend teams to understand source tables, relationships, and business logic
Help define repeatable patterns for onboarding new products and data sources
Target Platform
AWS
Databricks
Spark / PySpark
S3
Delta Lake
Unity Catalog
Databricks Jobs / Workflows
PostgreSQL
Python
SQL
Terraform
Git and CI/CD
What we’re looking for
Strong production experience in Data Engineering
Advanced Python and SQL skills
Hands-on production experience with Databricks and Spark/PySpark
Practical AWS experience, particularly with S3 and IAM
Experience ingesting data from PostgreSQL or other relational databases
Strong understanding of incremental pipelines, historical backfills, idempotency, retries, and reprocessing
Experience designing analytical data models and working with medallion architecture
Experience implementing data-quality checks, monitoring, reconciliation, and troubleshooting
Experience with Git-based development and CI/CD workflows
Ability to take ownership of complex data initiatives from design through production operation
Comfort working in a greenfield environment where standards, ingestion patterns, and models are still being defined
Ability to collaborate effectively with DevOps, Backend Engineering, Product, Analytics, and business stakeholders
Strong advantages
Experience with Terraform or another Infrastructure as Code tool
Hands-on experience with Unity Catalog
Experience with Databricks Jobs, Workflows, or Lakeflow
Understanding of AWS networking, VPCs, and EKS environments
Experience with CDC technologies such as AWS DMS or Debezium
Experience working with subscription and payments data
Familiarity with Stripe, Adyen, Solidgate, or other payment service providers
Experience integrating marketing or attribution data
Experience with dbt
Previous responsibility for defining data-platform standards or reusable engineering patterns
Experience building a data platform in a startup, scale-up, or other ambiguous environment
What success could look like
During your first stage in the role, you will help:
Establish the core AWS, Databricks, S3, Unity Catalog, and Terraform platform foundation
Productionize the first reusable end-to-end ingestion pattern
Onboard and unify payments data across multiple products
Deliver core Gold models for revenue, subscriptions, churn, and LTV
Create and document a repeatable approach for onboarding additional products
Put monitoring, reconciliation, CI/CD, and cost controls into production
What our client offers
Competitive compensation
Fully remote work with flexible working hours
22 paid vacation days plus local public holidays
A modern engineering environment with contemporary technologies
The opportunity to shape a growing Data function and its technical foundations
Meaningful platform challenges with room to influence architecture and engineering practices
A collaborative, product-focused environment where data directly supports business decision