Lead DevOps Engineer (12 Month FTC)
- london, england, United Kingdom
- Contract
- £85 - £110 Per Hour
Job Description
Lead DevOps Engineer - 12 Month FTC
Who We Are: AND Digital are a tech company focused on accelerating digital delivery and dedicated to closing the digital skills gap. We’ve been helping organisations build better digital products and stronger digital teams since 2014.
We believe our work should always leave a legacy for the client. We do this through close relationships with our offices (or ‘Clubs’) so that our partners are always prioritised by a regional team close to them.
This unique model has driven success for our clients and ourselves, evidenced by our remarkable organic growth since 2014. Today we number more than 1,300 people with Clubs all over the UK and Europe with plans for global expansion in the next couple of years.
Join us - and help us fulfil our mission to close the world’s digital skills gap.
The Role
We are looking for an experienced Lead DevOps Engineer to take technical ownership of the cloud infrastructure, deployment platforms and operational tooling that underpin our clients Accounts platform.
You will play a key leadership role in building and evolving a secure, resilient and highly scalable AWS environment, while establishing engineering standards and best practices across infrastructure, automation, observability and operational excellence.
Working closely with software engineering, security and platform teams, you will lead the design and implementation of reliable delivery processes and help ensure the Accounts platform can operate securely and efficiently at scale.
Key Responsibilities
- Provide technical leadership and ownership of the AWS cloud infrastructure supporting the Accounts platform.
- Design, build and continuously improve scalable, resilient and secure cloud architectures.
- Lead the development and maintenance of CI/CD pipelines, enabling reliable, automated and efficient software delivery.
- Establish and champion Infrastructure as Code practices, primarily using Terraform.
- Define and improve monitoring, observability, alerting and operational controls across the platform.
- Drive best practices around cloud security, networking, access control and infrastructure governance.
- Lead the technical response to production incidents, ensuring effective diagnosis, remediation and preventative action.
- Identify opportunities to improve platform reliability, performance, scalability and operational efficiency.


