Open roleExternal
24/7 HPC Infra SRE for AI & GPU Compute
- gloucester, england, United Kingdom
- Permanent·On-site
- Full time
- £90 - £120 Per Hour
Job Description
Radiant is seeking a senior Infrastructure Site Reliability Engineer to own and improve large‑scale GPU‑accelerated HPC infrastructure in a 24/7 production environment.
You will work across network, storage, virtualization and orchestration with hands‑on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking.
This role champions observability, automation and on‑call reliability, shaping next‑gen HPC platforms within a globally distributed team.
#J-18808-Ljbffr

