I
Posted 1mo ago•Remote
Principal Software Engineer - Data Platform (Iceberg/Trino)
principalRemoteSalary undisclosed
Required Skills
PythonJavaKubernetesSQLSnowflakeApache Spark
Job Description
About the Role
As a Principal Software Engineer on the Data Platform, you will own the architecture of the lakehouse engine that powers Innovaccer's platform in on-premise deployments: Apache Iceberg as the table format, Trino as the query engine, a REST catalog service, and Spark as transform compute. Cloud data warehouses have no on-premise equivalent, so this is a ground-up engine design, not a re-point. It is the single longest-lead technical track in the program, and the decisions you make on catalog, engine placement, and pipeline redesign gate everything downstream: transforms, serving, and reporting.
A Day in the Life
- Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout.
- Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores.
- Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators).
- Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup.
- Lead the SQL dialect strategy for porting existing warehouse
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.