Site Reliability Engineer - Vice President
About the Role
iCapital is looking to hire a Site Reliability Engineer to join the Site Reliability Engineering team, which plays a critical role in ensuring the platform delivers consistent, reliable service to clients. The services run on Kubernetes in AWS, and this role is focused on maintaining the health, stability, and performance of production environments.
This position is ideal for a hands-on engineer who thrives in diagnosing complex production issues where the root cause is not immediately apparent. This role transforms insights gained from production incidents into stronger reliability standards, safer deployment practices, more effective alerting, and automation that helps prevent recurring issues. While metrics, logs, and traces are leveraged daily, the core focus is identifying, resolving, and learning from the issues they reveal.
Responsibilities
- Act as a senior escalation point for complex production issues in Kubernetes-based services, leading hands-on investigation across application, container, Kubernetes, and AWS layers through to root cause.
- Partner with the teams that own our cluster in