Hardware Diagnostics Engineer - Infrastructure
About TensorWave
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
About the Role
We are looking for a Hardware Diagnostics Engineer to run burn-in, triage what fails, work servers out-of-band, and own RMAs end to end. If you like hardware that misbehaves in ways that take real work to explain, this is a good seat.
Before any GPU server carries a customer workload, it has to prove it works — under load, at temperature, for hours. When it doesn't, somebody has to figure out why, get replacement hardware in, and send the failed part back to the vendor.
What You’ll Do
Run server and GPU burn-in and stress testing, interpret the results, and decide whether hardware is production-ready
Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling: reproduce the failure, isolate the faulty component, and document what proved it