Open roleExternal
Frontier AI ML Engineer — Coding Agent Evaluator
- , , united kingdom, United Kingdom
- Job type not listed
- £203 - £407 Per Hour
Job Description
ML Engineer (Coding Agent Experience) ($85/hr)
- Contributors help evaluate and improve frontier AI coding models through structured technical assessments
- The work focuses on realistic machine learning engineering workflows and model evaluation
- Spots are limited and filling quickly on a first come, first serve basis
About The Role
- Contributors help evaluate and improve frontier AI coding models through structured technical assessments
- The work focuses on realistic machine learning engineering workflows and model evaluation
- Spots are limited and filling quickly on a first come, first serve basis
- Use frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks
- Review model-generated implementations involving model training, inference systems, MLOps, and LLM applications
- Identify bugs, edge cases, performance issues, and failure modes
- Compare outputs from multiple frontier models and assess their strengths and weaknesses
- Apply professional engineering judgment to realistic ML engineering scenarios
- Sprint based project that runs in 12-24 hour stretches based on client requirement
- 2+ years of professional machine learning engineering experience
- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products
- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools
- Ability to evaluate model-generated machine learning implementations and technical tradeoffs
- Experience deploying ML systems to production is preferred
- We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request
- $400 per accepted task
- Typical tasks take approximately 2-3 hours after ramp-up
- Compensation is tied to accepted work


