Open roleExternal
Head of Site Reliability Engineering – SRE
- bristol, england, United Kingdom
- £110 - £140 Per Hour
Job Description
- Define, lead, and evolve our global reliability strategy.
- Drive operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape.
- Work closely with Engineering, Infrastructure, Security, and Technology Operations teams to establish and embed modern SRE practices that enable highly reliable, scalable, and resilient services while fostering a culture of shared ownership and continuous learning.
- Drive adoption of SRE principles (SLOs, error budgets, toil reduction).
- Establish observability and monitoring standards.
- Lead automation-first operations.
- Improve incident and problem management maturity.
- Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.
- Establish SRE governance, standards, and operating model.
Requirements
- Proven experience building, leading, and developing Site Reliability Engineering or Production Engineering teams, with a strong understanding of SRE principles including service level objectives (SLOs), service level indicators (SLIs), error budgets, and toil reduction.
- Extensive experience driving automation initiatives, supported by strong scripting and development capabilities using technologies such as Python, PowerShell, Bash, Terraform and Ansible Automation Platform (AAP).
- Robust knowledge of observability and monitoring practices, with hands-on experience implementing and managing platforms such as Dynatrace, Prometheus, Grafana, and Splunk.
- Good understanding of CI/CD tooling and modern software delivery practices, including Jenkins, GitLab CI, and Azure DevOps.
- A background spanning both software engineering and technology operations environments.
- Professional certifications in cloud technologies, Site Reliability Engineering, platform engineering, or reliability engineering disciplines.
- Passionate about reliability, resilience, automation, and continuous improvement.
- A strategic thinker who can balance long-term vision with operational delivery.
Demonstrates expertise in Site Reliability Engineering principles, including SLOs, SLIs, and error budgets, while driving automation and observability initiatives. Proven ability to lead cross-functional teams in establishing reliable and resilient technology operations.
Highest-signal resume keywords
- Site Reliability Engineering Leadership
- Automation Initiatives
- Observability and Monitoring Practices
- CI/CD Tooling Experience
- Cloud Technology Certifications
Hard Skills
- Site Reliability Engineering Principles
- SLOs
- SLIs
- Error Budgets
- Automation Scripting
- Python
- PowerShell
- Bash
- Terraform
- Ansible Automation Platform
Soft Skills
- Strategic Thinking
- Passion for Continuous Improvement
Certifications & Qualifications
- Cloud Technologies
- Site Reliability Engineering
- Platform Engineering
- Reliability Engineering
Industry Keywords
- Operational Excellence
- Service Reliability
- Continuous Improvement
- Incident Management
- Problem Management
Tools & Technologies
- Dynatrace
- Prometheus
- Grafana
- Splunk
- Jenkins
- GitLab CI
- Azure DevOps


