Senior Software Engineer, Infrastructure
About the Role
This is a backend-architecture-heavy platform engineering role focused on owning the reliability, scale, performance, and developer experience of core infrastructure systems. You will join an engineering team of roughly 15 people working on an AI/ML evaluation and reinforcement learning platform. Your work directly shapes how fast, reliable, and cost-effective the platform is to build on and operate.
What You'll Do
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Wr
