SRE Architect
We fuse together exceptional talent who deliver outstanding software solutions. Our approach has helped us grow 60% in 2021, 94% in 2022, while in 2023 we joined forces with Insight, a Fortune 500 company and a leading solutions and systems integrator. With exciting growth plans and cutting-edge projects, there has never been a better time to join our incredible team.
SRE Architect / Principal Engineer
About the role
We are looking for a highly experienced SRE Architect / Principal Engineer to help lead the infrastructure and reliability architecture for the technology landscape as we accelerate our modernisation journey.
Our broader architectural direction is already taking shape, including DDD, microfrontends and Backend-for-Frontend (BFF), while our infrastructure direction is standardising around AWS, Terraform, GitHub Actions and ECS/Fargate. However, important decisions remain around areas such as gateways and routing, scalability, observability, resilience, deployment architecture, and how existing systems should progressively move toward the target state.
Your primary responsibility will be to consolidate this direction into a coherent infrastructure and reliability architecture and define a pragmatic path for its adoption.
You will operate between Enterprise Architecture and SRE/engineering teams, translating broader architectural direction into practical patterns, standards, reference implementations, and modernisation strategies.
This is a hands on Principal level individual contributor role with significant technical influence across infrastructure. You will be expected to challenge existing decisions where appropriate, validate important architectural choices through proofs of concept and reference implementations, and provide the technical direction that enables engineering teams to implement and adopt the target architecture successfully.
What you'll work on
Infrastructure modernisation & target architecture
- Define and evolve the target infrastructure and reliability architecture.
- Consolidate architectural decisions already underway into a coherent, scalable, and maintainable target state.
- Define pragmatic modernisation strategies for existing systems, balancing business value, technical risk, cost, and migration effort.
- Assess systems and recommend whether they should be incrementally modernised, aligned with the target architecture, temporarily retained, or eventually replaced.
- Define transition patterns that allow teams to modernise without unnecessary large scale rewrites.
- Establish the target architecture and infrastructure patterns as the default for new modules.
- Identify architectural gaps, risks, and cross-system dependencies across landscape.
AWS cloud & platform architecture
- Define scalable, resilient, secure, and cost-conscious architectures using AWS.
- Define approaches to service-to-service communication, external API exposure, routing, and gateway strategy.
- Guide architectural decisions around scalability, workload and capacity management, fault isolation, availability, and graceful degradation.
- Define appropriate environment, networking, and deployment strategies for services.
- Work closely with security and enterprise platform teams to ensure alignment with Pearson-wide standards.
Infrastructure as Code & CI/CD
- Establish Infrastructure as Code standards using Terraform, including reusable patterns, modules, and conventions that engineering teams can adopt consistently.
- Shape CI/CD architecture using GitHub and GitHub Actions as the standard delivery platform.
- Define reusable deployment patterns and approaches to environment promotion, rollback, and safe releases.
- Reduce infrastructure and CI/CD divergence between engineering teams through reusable standards and automation.
Reliability, observability & production readiness
- Shape and evolve wide approaches to reliability, resilience, observability, and production readiness.
- Shape standards for metrics, logs, traces, dashboards, and alerting across distributed systems.
- Help establish meaningful SLIs, SLOs, and reliability targets where appropriate.
- Guide architectural approaches to disaster recovery, failure handling, backups, and recovery strategies.