Engineering Manager SRE GPU
WHY WE NEED YOU ?
As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI and HPC infrastructure.
Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.
This is not a traditional IT Operations management role. You will combine engineering leadership with strong technical ownership, working close to the infrastructure itself from Linux systems, networking and bare-metal server provisioning to hardware lifecycle, automation and cluster observability.
You will help the team automate critical infrastructure workflows, improve reliability and operate production-grade GPU platforms powering our sovereign cloud.
YOUR FUTURE TEAM
We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.
You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.
The team works on some of our most critical AI and HPC infrastructure challenges, including bare-metal provisioning, GPU cluster automation, server lifecycle management, hardware failure management, observability, reliability and the integration of new GPU technologies.
The scope goes beyond traditional cloud-native infrastructure: the team operates close to the physical servers and needs to automate the full lifecycle of large fleets of GPU machines, from remote provisioning to production operations and remediation.
You will collaborate closely with GPU Cloud Engineering, Hardware, Product and Operations teams, as well as other infrastructure teams across Scaleway.
YOUR DAILY ROUTINE
Tasks
- Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
- Provide technical leadership and challenge architecture decisions related to large-scale GPU infrastructure
- Design and drive automation for bare-metal server provisioning and lifecycle management
- Improve remote deployment and server management capabilities using technologies and concepts such as PXE, BMC and IPMI
- Drive automation around hardware failures, remediation and server recovery
- Build and improve observability, monitoring, logging and alerting capabilities across production GPU clusters
- Own the SRE team's technical roadmap, priorities and delivery
- Ensure the reliability, scalability, performance and resilience of production GPU infrastructure
- Drive continuous improvements in automation, incident response and post-incident remediation
- Support technical decisions involving Linux systems, networking, hardware and cluster architecture
- Collaborate closely with Engineering, Hardware, Product, Operations and other GPU