Head of Engineering - GPU Cloud
WHY WE NEED YOU ?
As our GPU Cloud business continues to scale, we are strengthening our engineering leadership to support the deployment and operation of increasingly large and complex GPU clusters.
Your mission will be to lead our Support Engineering, HPC, and SRE teams, own key technology and architecture decisions, and ensure that our most strategic GPU infrastructure projects are successfully designed, delivered, and operated.
You will also play a critical role in validating the technical feasibility of commercial proposals and ensuring that the commitments we make to customers can be delivered reliably on our sovereign cloud infrastructure.
YOUR FUTURE TEAM
We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.
You will lead an organization of 14 engineers across two squads, each managed by an Engineering Manager reporting directly to you.
As part of the broader GPU Cloud organization, reporting to the SVP GPU Cloud, you will work closely with GTM, Operations, Service Management, Product, and other engineering teams to build and operate large-scale AI and HPC infrastructure.
YOUR DAILY ROUTINE
Tasks
- Lead the Support Engineering, HPC, and SRE organizations, directly managing two Engineering Managers responsible for 14 engineers
- Own the technical strategy, architecture, and key technology choices for GPU Cloud infrastructure
- Review and validate the technical and service dimensions of strategic commercial proposals, ensuring commitments are realistic and deliverable
- Oversee the design, deployment, and operational readiness of new GPU clusters
- Drive the evolution of our cluster management, capacity management, automation, and operational capabilities
- Ensure the reliability, scalability, performance, and maintainability of our GPU infrastructure
- Provide technical leadership on complex AI and HPC infrastructure projects
- Build strong alignment between Engineering, GTM, Product, and Operations
- Develop the engineering organization through clear direction, effective delegation, coaching, and long-term team development
- Establish and maintain high standards of engineering rigor, operational excellence, and technical decision-making
ABOUT YOU
HARDSKILLS:
- 10+ years of experience in infrastructure engineering, including significant experience leading senior technical teams and managers
- Proven experience designing, deploying, or operating large-scale infrastructure and compute clusters
- Familiarity in GPU and/or HPC environments, ideally involving NVIDIA and AMD technologies
- Strong understanding of distributed infrastructure, cluster architecture, reliability, and production operations
- Experience with orchestration, provisioning, and observability technologies such as Kubernetes, Proxmox, Wa