Senior Site Reliability Engineer
As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.
Responsibilities
Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
Lead production debugging and incident response: build and maintain runbooks, participate in the