Research Engineer, Benchmarks
HAMRA tailors your résumé to this role at Clera, writes a matching cover letter, and fills the application — you just review and press submit.
About the Role
This role sits at the heart of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to ensuring benchmarks are rigorous, credible, and practically meaningful.
What You'll Do
• Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
• Partner with subject-matter experts to define realistic workflows and translate them into evaluation criteria.
• Build reliable infrastructure to run models and agents against benchmark tasks at scale using Python, Docker, and Linux environments.
• Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes.
• Validate that benchmark performance correlates with real-world evaluations and customer needs.
• Write clear technical documentation and benchmark reports for research and engineering audiences.
What We're Looking For
• 2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
• Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.
• Demonstrated experience designing and running benchmarks or evaluation environments for AI agents or large language models.
• Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation.
• Experience collaborating with domain experts to translate workflows into structured evaluation tasks.
• Strong technical writing skills, with published papers or technical posts on AI benchmarking, model evaluation, or failure modes being a plus.
• Ability to reason from first principles about task design, scoring, and edge cases.
• Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.
• Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a bonus.
Compensation & Benefits
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore. This is a full-time, in-person role.
More roles to explore
- Founding Engineer (Full-Stack/AI)Clera · Munich
- Lead Product DesignerClera · Vienna
- Product DesignerClera · Paris
- Software Engineer, AI & Data SystemsClera · San Francisco
- Founders AssociateClera · Munich
- Forward Deployed Engineer, DACHClera · Munich
Listing sourced from Clera’s ashby board. HAMRA is not the employer; applications are completed on the employer’s own site.