Senior Site Reliability Engineer -AI Infrastructure Operations
About Nscale
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a senior SRE role for someone who sets the reliability bar and then pulls the rest of the team up
to it. You'll own the hardest problems on the platform: the automation other engineers build on, the
services that can't go down, and the design decisions that determine whether either holds up at scale.
You'll still carry a pager, but the real job is making sure it fires less, for everyone, over time.
What You'll Do
• Own reliability for critical production services end to end; set the direction, not just respond to what
breaks.
• Grow the team, not just the systems; mentor other SREs through design review, pairing, and
incident debriefs, and hold the bar that pulls everyone up to it.
• Set the standards the rest of the team works to: the SLO framework, the incident process, and