Senior Manager, Cluster Engineering & Deployment
About TensorWave
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
About the Role
The Senior Manager, Cluster Engineering & Deployment owns and runs the machine that turns delivered racks into accepted clusters: network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation and burn-in (including RCCL/collective performance), and the acceptance gate into production. This is one of the most schedule-critical roles in the pillar cluster revenue starts when this team says a cluster is ready.
What You’ll Do
Own the cluster deployment playbook and drive its evolution: staged bring-up, automated config push, link/optics validation, cabling verification against L1 port maps, and fault triage during deployment windows.
Lead deployment engineering across concurrent cluster builds, through team