T
Posted 5mo ago•George Town, Penang, Malaysia
Site Reliability Engineer (SRE) - Ads / Monetization Platform
MiddleOn-site (George Town)Salary undisclosed
Required Skills
PythonJavaAWSGCPAzureKubernetesDockerTerraformSQLLinux
Job Description
Role Summary
As a Site Reliability Engineer (SRE), you will build and operate highly available, globally distributed advertising/monetization services. You will improve reliability, scalability, and operability through automation, observability, incident management, and sound engineering practices.
Key Responsibilities
- Own reliability across the service lifecycle: design reviews, capacity planning, launch, deployment, operations, and continuous improvement.
- Build and operate highly available services across multiple regions/data centers; improve resilience, latency, and scalability.
- Develop automation and tooling to reduce toil (deployment, remediation, runbooks, self-healing) using scripting and software engineering best practices.
- Define and implement SLOs/SLIs/SLAs; create dashboards and alerting to track service health (availability, latency, errors, saturation).
- Lead sustainable incident response: triage, mitigation, root-cause analysis (RCA), and blameless postmortems with actionable follow-ups.
- Collaborate with software engineering, security, and compliance stakeholders to meet data governance and regulatory requirements.
Requirements
Must-have Qualifications
- 3+ years of experience in SRE, DevOps, systems engineering, or production operations for large-scale services.
- Strong coding skills in one language: Python or Go or C++
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.