N
Posted 1w agoHouston; New York; San Francisco; Seattle

Principal Infrastructure Engineer, AI Cluster Performance & Validation

principalOn-site (San Francisco)Salary undisclosed
Required Skills
PythonKubernetesDockerTerraformLinuxPyTorch
Job Description

Overview

As a Principal Infrastructure Engineer, AI Cluster Performance & Validation, you will be a critical member of the AI Infrastructure Operations team, responsible for ensuring the acceptance, performance, and scalability of our cutting-edge AI and High-Performance Computing (HPC) environments. Leveraging software engineering and testing principles, you will focus on building and maintaining the control plane, tooling, and automation that supports performance and validation testing of large-scale AI clusters. Your work will directly translate into higher system availability, compute optimization and reduced operational costs.

Key Responsibilities

  • Own the technical definition of "healthy at scale." Set the architecture, roadmap, and acceptance criteria by which multi-thousand-GPU clusters are declared production-ready, and establish the performance bar (collective bandwidth, job goodput, model FLOPs utilization) that every cluster must clear before and after customer handover. Establish technology and product direction in collaboration with other tech leads, managers, and senior leadership.
  • Run and instrument real AI workloads as a diagnostic instrument. Stand up and execute distributed training and inference jobs — open-source and customer-representative models — across thousands of accelerator

Similar Openings in AI & Machine Learning

View all in category