N
Posted 1w ago•Houston; New York; San Francisco; Seattle
Principal Infrastructure Engineer, AI Cluster Performance & Validation
principalOn-site (San Francisco)Salary undisclosed
Required Skills
PythonKubernetesDockerTerraformLinuxPyTorch
Job Description
Overview
As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs.
Key Responsibilities
- Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
- Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerator
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.