I
Posted 1mo agoRemote

Principal Software Engineer - Data Platform (Iceberg/Trino)

principalRemoteSalary undisclosed
Required Skills
PythonJavaKubernetesSQLSnowflakeApache Spark
Job Description

About the Role

As a Principal Software Engineer on the Data Platform, you will own the architecture of the lakehouse engine that powers Innovaccer's platform in on-premise deployments: Apache Iceberg as the table format, Trino as the query engine, a REST catalog service, and Spark as transform compute. Cloud data warehouses have no on-premise equivalent, so this is a ground-up engine design, not a re-point. It is the single longest-lead technical track in the program, and the decisions you make on catalog, engine placement, and pipeline redesign gate everything downstream: transforms, serving, and reporting.

A Day in the Life

  • Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout.
  • Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores.
  • Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators).
  • Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup.
  • Lead the SQL dialect strategy for porting existing warehouse

Similar Openings in Backend

View all in category