Sr. Site Reliability Engineer
About the Role
We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.
Key Responsibilities
Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently
Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior
Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems
Architect and deploy cloud infrastructure on AWS or Azure; implement infrastr