Open roleExternal
Inference Systems Performance Engineer for AI Serving
- london, england, United Kingdom
- Job type not listed
- £90 - £130 Per Hour
Job Description
Adaption is seeking a senior ML systems engineer to own the cost and performance of our inference stack in a rapidly evolving environment. You will shape throughput, latency, and reliability by tuning caching, batching, and kernels while collaborating with the serving fleet engineers.
You will work with engines like vLLM, SGLang, and TensorRT-LLM, and build tools to measure where compute is spent. A strong background in Python and systems languages is essential, with a track record of real-world
#J-18808-Ljbffr

