Site Reliability Engineer
- charing cross, london, SW1A, United Kingdom
- Permanent·On-site
- Full time
Job Description
Scroll down to find the complete details of the job offer, including experience required and associated duties and tasks.
Site Reliability Engineer
Up to £85,000 + Benefits
Central London Hybrid (2/3 days a week in the office)
Build, Scale & Improve the Reliability of a Fast-Growing SaaS Platform
We're partnering with a fast-growing SaaS company that's going through an exciting period of growth and investing heavily in its engineering and platform capabilities.
They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure.
This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience.
You'll have genuine ownership and the opportunity to influence how the platform evolves as the business continues to scale.
What You'll Be Doing
- Design, build, and maintain highly available and scalable AWS infrastructure
- Manage and optimise Kubernetes environments and containerised workloads
- Build and maintain infrastructure using Terraform and Infrastructure as Code principles
- Develop and optimise CI/CD pipelines using GitHub Actions
- Build and improve comprehensive monitoring and observability across the platform
- Implement and maintain effective logging, metrics, tracing, alerting, and dashboards
- Define and improve SLIs, SLOs, and reliability metrics
- Proactively identify and resolve performance, availability, and reliability issues
- Lead and contribute to incident response, troubleshooting, and root cause analysis
- Automate operational processes and eliminate repetitive manual tasks
- Work closely with software engineers to improve deployment processes, system xwwtmva reliability, and developer experience
- Help improve platform resilience, scalability, and disaster recovery capabilities
- Contribute to capacity planning and performance optimisation as the platform scales
- Establish and champion SRE best practices across the wider engineering function
What We're Looking For
- Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer, or similar
- Strong hands-on experience with AWS
- Strong experience working with
Please click on the apply button to read the full job description


