AI Infrastructure Engineer
MiddleOn-site (Fremont)Salary undisclosed
Required Skills
AWSGCPAzureKubernetesDockerTerraform
Job Description
- Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads
- Own datacenter and infrastructure operations: compute, storage, networking, and the automation layered on top
- Build and maintain Infrastructure-as-Code using Terraform (Terragrunt), Ansible, etc.
- Deploy, secure, and operate our HashiCorp stack (Vault, Boundary) alongside identity/access via Keycloak
- Build and maintain observability with Prometheus, Grafana, and Alertmanager
- Deploy and manage containerized workloads via Docker and Kubernetes
- Design and maintain networking infrastructure: VLANs, routing, firewall rules, and load balancers
- Write scripts/tools (and contribute code where useful) to automate operational toil
- Provide general datacenter support for customers colocated in the HostMax datacenter — things like cabling, hardware swaps, rack/stack work, and basic troubleshooting
- Participate in business-hours (M-F) on-call rotation for customer and infrastructure support
- Participate in an SDLC/Scrum-based workflow, using Jira and Confluence for planning and documentation
- Help define standards, runbooks, and best practices as the team grows
- Development and deployment of AI infrastructure workloads on-prem (GPU scheduling, model serving infra, self-hosted AI tooling, etc.)
Requirements
- Solid experience
Ready to apply? Optimize your CV for this specific jobAI customizes your experience bullets and increases chances to get hired.