Cloud Systems Engineer
We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads.
The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure.
Responsibilities:
AI Infrastructure Operations
- Deploy, configure, and maintain GPU-accelerated compute infrastructure.
- Manage operating system, firmware, BIOS, BMC, driver, and software lifecycle updates.
- Monitor system health, performance, utilization, and capacity across AI infrastructure environments.
- Support infrastructure utilized for AI model training, inference, and data processing workloads.
- Develop and maintain operational standards, runbooks, and maintenance procedures.
- Participat