Site Reliability Engineer - Data Platform
IMC operates on the cutting-edge use of technology to create a competitive edge over the competition. We also grow quick and have plenty of complex technical challenges. We're looking for an experienced SRE with strong background in managing distributed data systems on both bare metal Linux and Kubernetes. We want someone who can help us standardize deployments, elevate observability, and improve automation as we scale our data platform and other critical data services.
You will join our Data Platform team, part of our local data team that builds and runs the systems that are used by traders, quant researchers and engineering teams, for all their data needs. The Data Platform team is responsible for the foundational platform that our data frameworks and tooling is built on top of. This includes observability, scalability and supporting standardised deployments.
Your Core Responsibilities:
As an SRE within IMC you will join a sub-team that takes a central role in all the data needs and you’ll be working to:
- Design, implement and operate our data platforms.
- Improve observability so we catch issues before our users do.
- Build automation to reduce toil and allow our systems to scale
- Support and own reliability of critical services (e.g. HDFS, Kafka and Dremio)
