Senior Lead Software Engineer - LLM Ops Platform Reliability
- glasgow, city of glasgow, G1, United Kingdom
- Permanent·On-site
- Full time
Job Description
Applying for this role is straight forward Scroll down and click on Apply to be considered for this position.
hackajob is partnering directly with JPMorganChase to hire for this role.
JOB DESCRIPTIONHelp shape how AI systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here.
As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the large language model inference platform end to end. You will operate large language model serving stacks in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability, and lead incident response and continuous improvement.
Job responsibilities
- Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure
- Build backend services and APIs that enable reliable operation of AI infrastructure in production environments
- Operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization
- Deploy, host, and lifecycle-manage open-source and proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible xwwtmva infrastructure as code and continuous delivery pipelines
- Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads
- Tune GPU and accelerator capacity, autoscaling, and cost efficiency for large language model inference workloads using performance optimization techniques such as quantization, parallelism, and speculative decoding
- Lead reliability engineering for large language model endpoints through capacity planning, load and soak testin
Please click on the apply button to read the full job description


