Principal Software Engineer - Fleet Management

principalOn-site (AMER)Salary undisclosed
Required Skills
PythonNode.jsAWSGCPKubernetesTerraform
Job Description

About Nscale

Nscale is building a vertically integrated GenAI cloud from sustainable data centers to advanced AI infrastructure and enterprise applications. Our culture values open collaboration, ownership, and excellence.

About the Role

We're hiring a Principal Software Engineer to lead the technical development of our Fleet Manager platform - the workflow automation system that provisions, tests, and remediates GPU nodes and network switches at scale.

As technical lead, you'll own the architecture and delivery of foundational Python-based automation systems that manage the entire lifecycle of our compute infrastructure: device enrolment, burn-in testing, network configuration, GPU health monitoring, and self-healing capabilities. You'll mentor a team of senior engineers, set technical direction, and drive engineering excellence while remaining hands-on with critical systems.

What you'll do

  • Lead technical architecture and roadmap for Fleet Manager's workflow automation systems
  • Own end-to-end delivery of device provisioning, validation, testing, and remediation workflows at scale
  • Design and build workflow orchestration systems for GPU node and network switch lifecycle management
  • Establish engineering standards for reliability, observability, and operational excellence across all Fleet Manager services
  • Mentor and raise the bar for a team of senior engineers through design reviews, technical leadership, and hands-on collaboration
  • Drive architecture decisions balancing automation complexity, reliability, and maintainability
  • Integrate with infrastructure tooling: DCIMs, NetBox, OpenStack, bare metal APIs (MAAS, Ironic, IPMI)
  • Partner with Infrastructure, Platform, and SRE teams to translate operational needs into robust, scalable automation
  • Build production-grade Python systems for hardware lifecycle automation, leveraging AI tools to accelerate delivery

About You

  • You have 12-15+ years software engineering experience building and operating production systems, with proven technical leadership in infrastructure automation or workflow tooling
  • Strong Python engineering fundamentals with experience leading complex, multi-service distributed systems
  • You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement
  • Technical expertise: quickly understanding systems design tradeoffs, keeping track of rapidly evolving software systems
  • Track record of owning technical roadmaps and delivering large-scale automation systems from ambiguous requirements to production
  • You use AI tools like Claude or Cursor as a core part of your development workflow - as a fundamental multiplier of what you can build
  • Deep understanding of operational excellence: SLOs, monitoring, alerting, incident response, and production reliability
  • Strong mentorship skills with ability to develop high-performing engineering teams
  • Excellent communication skills to build consensus with stakeholders, both internally and externally

Strong candidates will have

  • Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
  • Bare metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
  • Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
  • Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
  • Track record of 1+ years leading large-scale, complex projects or technical teams
  • Open-source contributions in infrastructure automation or cloud-native tooling

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range
$240,000—$400,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Similar Openings in Backend

View all in category➔