SpaceX logo
Posted 1h agoHawthorne, CA

Site Reliability Engineer (High Performance Computing)

MiddleOn-site (Hawthorne)Salary undisclosed
Required Skills
PythonNext.jsNode.jsKubernetesDockerTerraformLinuxPyTorchTensorFlow
Job Description

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING)

SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process.

We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You’ll work alongside HPC systems engineers who design and commission clu

Similar Openings in DevOps & Cloud

View all in category