Senior Data Engineer
AppsFlyer processes 150+ billion events across its platform, and our Multi-Platform Measurement team is at the center of extending that power beyond the mobile device. We build the measurement infrastructure behind two of the company’s most strategic frontiers: user-level, cross-platform LTV measurement — stitching a single user’s journey across mobile, CTV, console and PC — and attribution for extended platforms — smart TVs, consoles and native PC — where there is no advertising ID and the classic mobile attribution playbook simply doesn’t apply.
We’re looking for a Senior Data Engineer to own measurement domains end to end: from raw event ingestion through distributed batch and streaming processing into the serving layer that powers our clients’ dashboards and APIs. This isn’t internal tooling — you’ll be building the product surface that advertisers and marketers rely on to understand performance on screens where nobody else can measure it yet. You’ll work across a Scala/Spark + Airflow batch stack on GCP and a Clojure + Spark Structured Streaming stack on Kafka, DynamoDB and EMR, and you’ll be expected to have real influence on where both are heading.
As a senior member of the team you’ll set the technical bar: leading design for significant initiatives, raising the quality of what the team ships through review and mentorship, and being the person other teams come to when cross-platform measurement is on the line.
As AI reshapes how data is processed, queried, and served, you’ll also have the opportunity to integrate AI and LLM-based capabilities into our data workflows — from intelligent data validation and anomaly detection to AI-powered analytics interfaces that make our platform smarter for clients.
What you’ll do
- Own Measurement Domains End-to-End: Take architectural ownership of significant parts of our cross-platform and extended-platforms measurement stack — from ingestion and distributed processing (batch and streaming) through to the serving layer — with strict freshness and accuracy SLAs on billions of events.
- Design for Ambiguity: Lead the design of new measurement capabilities where the rules aren’t written yet — identity stitching without device IDs, cross-platform user resolution, attribution semantics for CTV and console. You’ll turn fuzzy product intent into a concrete, defensible data design and write the doc the rest of the team builds from.
- Ensure Data Integrity: Own the correctness of complex attribution logic — deduplication, organic vs. non-organic classification, primary/secondary attribution resolution, reattribution and reengagement flows, and cross-platform user stitching across massive, partitioned datasets.
- Performance Engineering: Optimize scalable distributed processes to meet strict client SLAs. Tune jobs, queries, and pipeline scheduling for maximum throughput and minimum latency — and know when the answer is a better data model rather than a bigger cluster.
- Ship Client-Facing Features: Partner with Product, R&D, and client-facing teams to turn raw signals from TVs, consoles and PCs into production-grade, queryable data products that advertisers interact with directly.
- Raise the Team’s Bar: Set standards through code review, design review, and hands-on mentorship. Onboard new engineers into the domain, write the playbooks, and make the systems you own understandable to people who didn’t build them.
- Operate What You Build: Own production health for your domains — monitoring, alerting, on-call, incident response, and the postmortems that prevent the next one.
- Leverage AI in Data Workflows: Explore and integrate AI/ML capabilities into the data platform — from automated data-quality checks and anomaly detection to LLM-powered analytics tools that enhance how clients interact with their data.
What you have
- 6+ years in Data / Big Data Engineering, with a track record of owning large-scale production systems — not just contributing to them.
- Deep experience with distributed processing frameworks (e.g., Spark) — you read execution plans and Spark UIs fluently, reason about partitioning, shuffles, skew and memory management, and can diagnose a performance regression from first principles.
- Production experience with streaming or near-real-time pipelines (e.g., Spark Structured Streaming, Kafka) — you understand state, watermarks, exactly-once semantics and what actually breaks at 3am.
- Solid SQL skills and experience with analytical data warehouses as both a processing and serving layer, including data modeling decisions for query performance and cost.
- Strong programming fundamentals — clean, testable, production-grade code, with the judgment to know what deserves abstraction and what doesn’t. Scala preferred.
- Comfortable being polyglot: our extended-platforms backend is functional (Clojure, some Go), and you’re the kind of engineer who reads an unfamiliar functional codebase and gets productive in it rather than rewriting it.
- Experience with workflow orchestration (e.g., Airflow) and cloud compute environments at production scale; comfortable across more than one cloud (we run on both GCP and AWS).
- A natural instinct for data quality: you think about edge cases, dedup logic, and “what happens when this field is null” before anyone asks — and you build the checks that catch it in production.
- Demonstrated technical leadership: you’ve led a multi-quarter initiative, mentored engineers, and driven a design through disagreement to a decision.
- Strong communication and stakeholder skills — you can explain a subtle attribution behavior to Product, and a subtle Spark behavior to a junior engineer, on the same day.
- B.Sc. in Computer Science or equivalent.
Bonus points
- Experience building or maintaining attribution / AdTech / MMP data systems.
- Familiarity with CTV, console, or connected-device measurement, and identity resolution without device-level advertising IDs.
- Advanced Scala or functional programming experience; hands-on Clojure a strong plus.
- Exposure to data modeling for multi-touch or last-click attribution, and to LTV / cohort measurement.
- Experience taking ownership of a system built by another team and successfully modernizing it.
- Hands-on experience with LLMs, ML pipelines, or AI-powered data tools.
- Experience with large-scale analytical warehouses at petabyte scale, and with cost optimization at that scale.
As a global company operating from 25 offices across 19 countries, we reflect the human mosaic of the diverse and multicultural world in which we live. We ensure equal opportunities for all of our employees and promote the recruitment of diverse talents to our global teams without consideration of race, gender, culture, or sexual orientation. We value and encourage curiosity, diversity, and innovation from all our employees, customers, and partners.