Manager, Site Reliability Engineer
MiddleOn-site (South San Francisco)Salary undisclosed
Required Skills
PythonNext.jsAWSGCPAzureKubernetesTerraformCI/CD
Job Description
Manager, Site Reliability Engineer (Hybrid in South San Francisco)
About the Role
We are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability, scalability, security, and performance of our cloud infrastructure and critical applications.
The ideal candidate combines strong technical expertise with leadership experience and a passion for operational excellence. You will lead efforts across cloud infrastructure, observability, automation, networking, and platform reliability while partnering closely with engineering and product teams to support business-critical applications.
Responsibilities
- Lead and manage the Site Operations / SRE function
- Own cloud infrastructure architecture, operations, scalability, and optimization
- Ensure high availability and reliability of production applications and services
- Drive operational excellence through automation, monitoring, and incident management
- Develop and maintain observability platforms including logging, metrics, alerting, and tracing
- Manage Kubernetes-based inf
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.