Senior Platform Engineer — AI Infrastructure
HelloPrint is mid-transformation. The entire platform is being rebuilt from the ground up: new frontend, new pricing engine, new content engine, new product engine. Everything agent-ready. We ship in a week what used to take a year. What we do not yet have is someone who owns the reliability of all of it: the SLOs, the cost, the AI runtime, and the guardrails that let product engineers move fast without breaking things. That is this role.
Core Objective: Take full technical ownership of production reliability, distributed observability, deployment safety, cost optimization, and AI runtime infrastructure across high-velocity microservices and cloud workloads.
What you will do:
SLOs & Error Budgets: Define, track, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error-budget policies across core customer journeys and critical services (including checkout, payments, catalog pipelines, and routing engines).
Distributed Observability & Telemetry: Expand Telemetry, Sentry + Google Cloud Monitoring distributed tracing, and automated diagnostic tooli