Scaleway logo
Posted Jul 16Paris
Apply ↗

Engineering Manager SRE GPU

MiddleOn-site (Paris)€60,000 – €80,000 / yr
Required Skills
Next.jsKubernetesLinuxPrometheusGrafana
Job Description
OUR STORY:
 
Join Scaleway and shape the sovereign cloud of tomorrow !
Since 1999, we have been designing secure, sustainable infrastructures aimed at supporting the most ambitious companies.
 
Historically known for our dedicated servers (Dedibox), we made a strategic shift to cloud computing in 2015. Staying true to our principles of simplicity, flexibility, and technical excellence, we have become one of the leading players in Europe in the sector.
 
With the rise of artificial intelligence, we have strengthened our commitment, supported by the Iliad Group, which is investing €3 billion to develop a serious, sovereign AI alternative to American and Asian giants.
 
Every day, thanks to our fast-growing portfolio of cloud and AI products (bare metal, containerization, serverless, AI, etc.), Scaleway proudly serves thousands of customer across the private and public sector, from corporations like France Télévisions or Hachette Livre, to fast-growing startups like Photoroom and Biolevate, to institutions like the City of Copenhagen.
 
Our offices are located in Paris, Lille, Toulouse, Rennes, Rouen, Bordeaux and Lyon.
 

WHY WE NEED YOU ? 

As our GPU Cloud infrastructure continues to scale, we are strengthening our SRE organization to support the deployment and operation of increasingly large and complex AI and HPC infrastructure.

Your mission will be to lead our Site Reliability Engineering team and ensure the reliability, scalability, and operational excellence of our GPU clusters.

This is not a traditional IT Operations management role. You will combine engineering leadership with strong technical ownership, working close to the infrastructure itself from Linux systems, networking and bare-metal server provisioning to hardware lifecycle, automation and cluster observability.

You will help the team automate critical infrastructure workflows, improve reliability and operate production-grade GPU platforms powering our sovereign cloud.

YOUR FUTURE TEAM 

We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.

You will lead a team of 6 Site Reliability Engineers within the GPU Cloud organization.

The team works on some of our most critical AI and HPC infrastructure challenges, including bare-metal provisioning, GPU cluster automation, server lifecycle management, hardware failure management, observability, reliability and the integration of new GPU technologies.

The scope goes beyond traditional cloud-native infrastructure: the team operates close to the physical servers and needs to automate the full lifecycle of large fleets of GPU machines, from remote provisioning to production operations and remediation.

You will collaborate closely with GPU Cloud Engineering, Hardware, Product and Operations teams, as well as other infrastructure teams across Scaleway.

 

YOUR DAILY ROUTINE 

Tasks

  • Lead and manage a team of 6 Site Reliability Engineers, supporting both their technical execution and career development
  • Provide technical leadership and challenge architecture decisions related to large-scale GPU infrastructure
  • Design and drive automation for bare-metal server provisioning and lifecycle management
  • Improve remote deployment and server management capabilities using technologies and concepts such as PXE, BMC and IPMI
  • Drive automation around hardware failures, remediation and server recovery
  • Build and improve observability, monitoring, logging and alerting capabilities across production GPU clusters
  • Own the SRE team's technical roadmap, priorities and delivery
  • Ensure the reliability, scalability, performance and resilience of production GPU infrastructure
  • Drive continuous improvements in automation, incident response and post-incident remediation
  • Support technical decisions involving Linux systems, networking, hardware and cluster architecture
  • Collaborate closely with Engineering, Hardware, Product, Operations and other GPU
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.

Similar Openings in DevOps & Cloud

View all in category