Site Reliability Engineer
MiddleHybrid (Hyderabad)Salary undisclosed
Required Skills
PythonAWSAzureKubernetesDockerTerraformCI/CD
Job Description
The Site Reliability Engineer is responsible for ensuring the availability, performance, and resilience of the organization's digital banking and financial services platforms. This role focuses on automating operational processes, defining and maintaining service-level objectives, and engineering systems that can withstand and recover from failure. You will work closely with engineering, DevOps, QA, cybersecurity, and compliance teams to ensure platform reliability meets both technical and regulatory standards, while minimizing risk to production systems through proactive monitoring, incident response, and continuous improvement of the software delivery lifecycle.
How You’ll Make an Impact:
Reliability Planning & Governance
- Define and maintain service-level objectives (SLOs), error budgets, and reliability targets aligned with business goals and compliance deadlines.
- Oversee the end-to-end service lifecycle, from code integration to production deployment, with a focus on stability and risk reduction.
- Ensure all changes comply with relevant financial regulations.
- Conduct reliability risk and blast-radius assessments before production changes.
- Coordinate go/no-go decisions with engineering, QA, compliance, and operation
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.