About the role

We are looking for a highly experienced SRE Architect / Principal Engineer to help lead the infrastructure and reliability architecture for the technology landscape as we accelerate our modernisation journey. Our broader architectural direction is already taking shape, including DDD, microfrontends and Backend-for-Frontend (BFF), while our infrastructure direction is standardising around AWS, Terraform, GitHub Actions and ECS/Fargate. However, important decisions remain around areas such as gateways and routing, scalability, observability, resilience, deployment architecture, and how existing systems should progressively move toward the target state. Your primary responsibility will be to consolidate this direction into a coherent infrastructure and reliability architecture and define a pragmatic path for its adoption. You will operate between Enterprise Architecture and SRE/engineering teams, translating broader architectural direction into practical patterns, standards, reference implementations, and modernisation strategies. This is a hands on Principal level individual contributor role with significant technical influence across infrastructure. You will be expected to challenge existing decisions where appropriate, validate important architectural choices through proofs of concept and reference implementations, and provide the technical direction that enables engineering teams to implement and adopt the target architecture successfully.

What you'll work on

  • Infrastructure modernisation & target architecture: Define and evolve the target infrastructure and reliability architecture; consolidate architectural decisions; define pragmatic modernisation strategies; assess systems; define transition patterns; establish target architecture as default; identify architectural gaps.
  • AWS cloud & platform architecture: Define scalable, resilient, secure, and cost-conscious architectures; define approaches to service-to-service communication; guide architectural decisions around scalability and availability; define deployment strategies.
  • Infrastructure as Code & CI/CD: Establish Infrastructure as Code standards using Terraform; shape CI/CD architecture using GitHub and GitHub Actions; define reusable deployment patterns; reduce infrastructure and CI/CD divergence.
  • Reliability, observability & production readiness: Shape approaches to reliability, resilience, and observability; shape standards for metrics, logs, and traces; establish SLIs and SLOs; guide architectural approaches to disaster recovery.
  • Architecture standards & AI-enabled engineering: Translate decisions into reusable standards and reference implementations; collaborate on AI-enabled engineering frameworks; create architecture decision records.
  • Technical leadership: Act as a senior technical authority; operate as a bridge between Enterprise Architects and engineering teams; validate decisions through proofs of concept; review major designs; mentor senior engineers.

We fuse together exceptional talent who deliver outstanding software solutions. Our approach has helped us grow 60% in 2021, 94% in 2022, while in 2023 we joined forces with Insight, a Fortune 500 company and a leading solutions and systems integrator. With exciting growth plans and cutting-edge projects, there has never been a better time to join our incredible team.

Architecture & technical leadership

  • Extensive professional experience designing and operating large-scale distributed systems in production.
  • Proven experience operating at SRE Architect / Principal Engineer or equivalent senior technical leadership level.
  • Strong track record defining cloud and infrastructure architecture across multiple teams or services.
  • Experience leading or shaping modernisation across technology estates containing both legacy and modern systems.
  • Ability to define a target architecture while creating realistic incremental migration paths toward it.
  • Strong architectural judgment and the ability to balance technical quality, delivery speed, risk, cost, and organisational constraints.
  • Strong technical depth and willingness to create prototypes or reference implementations.

AWS, infrastructure & delivery

  • Deep expertise in AWS and cloud-native architecture.
  • Strong experience designing and operating containerised workloads on AWS, particularly ECS/Fargate.
  • Strong understanding of AWS networking, load balancing, routing, and secure connectivity.
  • Experience designing scalable API, gateway, and ingress architectures.
  • Strong experience with relational data platforms such as Amazon RDS/Aurora.
  • Good understanding of event-driven and asynchronous architecture patterns.
  • Strong experience with Terraform in production environments.
  • Strong experience with GitHub and GitHub Actions, including reusable CI/CD workflows and deployment automation.
  • Strong understanding of cloud native autoscaling, workload capacity management, and scaling strategies.

SRE & reliability engineering

  • Strong understanding of Site Reliability Engineering principles.
  • Experience designing observability approaches for distributed production systems.
  • Strong understanding of metrics, logging, tracing, alerting, and production monitoring.
  • Experience working with SLIs, SLOs, reliability targets, and production-readiness practices.
  • Strong knowledge of resilience patterns, failure modes, scalability, and disaster recovery.
  • Ability to reason about trade-offs between reliability, performance, complexity, and cost.

Security

  • Strong understanding of AWS security fundamentals, including IAM, secrets management, encryption, and secure infrastructure patterns.
  • Ability to work effectively with specialist security teams.

General & soft skills

  • Strong technical leadership skills without relying on formal line-management authority.
  • Ability to influence engineering direction across multiple teams and organisational boundaries.
  • Comfortable challenging existing approaches constructively.
  • Pragmatic approach to modernisation.
  • Excellent communication skills with engineers, Tech Leads, Engineering Managers, Enterprise Architects, and other stakeholders.
  • Ability to turn ambiguous technical problems into clear decisions and actionable next steps.

Nice to have

  • Experience with Backend-for-Frontend and microfrontend architectures.
  • Experience with event-driven architectures and Kafka/MSK.
  • Experience with OpenTelemetry and modern observability platforms.
  • Experience with progressive delivery approaches such as canary or blue/green deployments.
  • Experience with cloud cost optimisation or FinOps practices.
  • Experience defining architecture standards, paved roads, or engineering golden paths across multiple autonomous teams.
  • Experience working in EdTech, digital learning, or another large global technology organisation.