Site Reliability Engineer, AI Observability
About us
Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.
We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.
Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.
Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.
ROLE RESPONSIBILITIES
- Platform Operations: Diagnose and resolve failures across ClickHouse