Senior Site Reliability Engineer
Job Summary
The Senior Site Reliability Engineer plays a vital role in ensuring the reliability, availability, and performance of DeepHealth software applications, which integrate AI algorithms to deliver clinically relevant information for enhanced decision support. This role takes ownership of the reliability of cloud components deployed at client sites, ensuring that the DeepHealth solution is scalable, resilient, and secure, providing support to the operational team, and providing technical leadership within the platform engineering practice.
Essential Duties and Responsibilities
Own the reliability, availability, and performance of cloud components deployed at client sites and of the DeepHealth solution.
Define and monitor service level objectives (SLOs), error budgets, and key reliability metrics.
Develop and implement automation tools and processes to eliminate toil and streamline deployment, monitoring, and incident response operations.
Design and maintain observability tooling (monitoring, logging, alerting, and tracing), and resolve issues before they impact clients.
Contribute to the writing of technical specifications and documentation, ensuring compliance with regulatory requirements and industry best practices.
Lead inci