Cloud Operations - Service Reliability Engineer
- BT39 9TU
- Permanent
- Full time
Job Description
What you will do
The Service Reliability Engineer is accountable for improving the reliability, observability and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering and automation, using tools and practices such as Bicep, Azure DevOps, GitHub and Ansible to support consistent, repeatable and well-governed service operation.
The role involves:Monitoring, observability and alerting across cloud infrastructure, platform services and supported application environments;
Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery; and
Promoting service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
Improve the reliability, resilience and performance of cloud-hosted services through monitoring, observability and automation.
Design and enhance monitoring, alerting and operational dashboards to provide real-time insight into service health and performance.
Support and develop Azure cloud platforms, working across infrastructure, networking, identity and platform services.
Implement Infrastructure as Code and automation solutions using tools such as Bicep, Azure DevOps, GitHub and Ansible.
Investigate incidents, identify root causes and drive continuous service improvements through automation and operational excellence.
Collaborate with engineering, infrastructure and support teams to improve operational standards, service supportability and platform resilience.
Provide technical expertise in cloud engineering, reliability engineering and observability best practices.
Contribute to the delivery of scalable, secure and highly available cloud services across a global technology environment.
Promote modern engineering practices, including source control, CI/CD pipelines, Infrastructure as Code and automated configuration management.
Create and maintain operational documentation, support guides and technical standards to enable effective service support.
What you will have
Business Competencies
Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
Technically curious, with an enthusiasm for understanding a broad set of systems, technologies and operational domains.
Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
Ability to make sound decisions under pressure and support effective incident response.
Strong commitment to service reliability, operational resilience and excellent customer service.
Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.
Excellent communication skills, both orally and written.
Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation and cloud engineering standards.
Experience of working in a global environment across international locations with an appreciation of multiple cultures.
Knowledge
Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.
Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.
Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.
Knowledge of configuration management and automation tooling such as Ansible.
Expected to develop a broad understanding of the technologies, systems and business working practices used by A&O Shearman.
Experience
Minimum 4–5 years’ IT experience with at least 2 years’ experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.
Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.
Experience of using or implementing monitoring using Elastic is desirable.
Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.
Experience working with diverse internal teams to improve service supportability, resilience and operational standards.
Experience of working in an ITIL environment.
NO AGENCIES PLEASE - A&O Shearman does not accept unsolicited CVs. For further information, please see our UK Recruitment Agency Policy and our commitment to direct sourcing .
#LI Job Code - Manager


