Senior Site Reliability Engineer
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฏ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฎ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฏ-๐ฎ๐ฌ ๐๐ฃ๐)
Experience: 4+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for an experienced Senior Site Reliability Engineer (SRE) to build, operate, and continuously improve highly reliable, scalable, secure, and high-performing production systems across hybrid and multi-cloud environments.
The role combines cloud infrastructure, Kubernetes, automation, observability, incident management, and reliability engineering. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, Python, Bash, and modern observability platforms, along with a strong understanding of production operations and distributed systems.
Requirements
Key Responsibilities
- Define and manage SLIs, SLOs, SLAs, error budgets, and reliability objectives for critical production services.
- Drive initiatives to improve system availability, scalability, performance, resilience, and operational efficiency.
- Manage and support production Kubernetes environments, including Amazon EKS and Red Hat OpenShift.
- Deploy and maintain containerised workloads using Docker, Kubernetes, and Helm.
- Manage cloud infrastruct