AMAX logo
Posted 2mo ago•Fremont, California, United States

AI Infrastructure Engineer

MiddleOn-site (Fremont)Salary undisclosed
Required Skills
AWSGCPAzureKubernetesDockerTerraform
Job Description
  • Design, build, and operate on-prem infrastructure that behaves like a cloud environment for internal teams, including AI/ML workloads
  • Own datacenter and infrastructure operations: compute, storage, networking, and the automation layered on top
  • Build and maintain Infrastructure-as-Code using Terraform (Terragrunt), Ansible, etc.
  • Deploy, secure, and operate our HashiCorp stack (Vault, Boundary) alongside identity/access via Keycloak
  • Build and maintain observability with Prometheus, Grafana, and Alertmanager
  • Deploy and manage containerized workloads via Docker and Kubernetes
  • Design and maintain networking infrastructure: VLANs, routing, firewall rules, and load balancers
  • Write scripts/tools (and contribute code where useful) to automate operational toil
  • Provide general datacenter support for customers colocated in the HostMax datacenter — things like cabling, hardware swaps, rack/stack work, and basic troubleshooting
  • Participate in business-hours (M-F) on-call rotation for customer and infrastructure support
  • Participate in an SDLC/Scrum-based workflow, using Jira and Confluence for planning and documentation
  • Help define standards, runbooks, and best practices as the team grows
  • Development and deployment of AI infrastructure workloads on-prem (GPU scheduling, model serving infra, self-hosted AI tooling, etc.)

Requirements

  • Solid experience

Similar Openings in AI & Machine Learning

View all in category➔