HRKyle Services · CEO
Experience: Technical aptitude and learning potential matter more than years
Pay: $100k-$200k, full time
Location: San Francisco, CA, US or Singapore (on-site)
Visa: Relocation and visa support available for strong candidates
The role
Build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You will create benchmarks that are technically rigorous, practically useful, and credible to frontier labs.
What you'll do
Own the design, implementation, and quality of HUD's internal agent benchmarks
Work with subject-matter experts to define realistic domain-specific tasks
Build infrastructure to run models and agents reliably against benchmark tasks
Develop metrics and analyses for benchmark difficulty, reliability, and failure modes
Validate whether benchmark performance correlates with real-world evaluations, customer needs, and lab expectations
Write clear documentation and benchmark reports for technical audiences
What we're looking for
Proficiency in Python, Docker, and Linux environments
Published papers or technical blogs about benchmarks, model failure modes, or related topics
A strong understanding of what makes a benchmark realistic, reliable, and useful
Experience with environments and evaluations
Curiosity and the ability to understand how workflows in unfamiliar domains actually work
First-principles reasoning about task design, scoring, and failure modes
Strong signals
Detail orientation and an eye for subtle inconsistencies and edge cases
Independent work in unstructured problem spaces
Early-stage startup experience
Strong communication across time zones
Data Analysis
Experimental Design
Statistical Modeling
Technical Reporting
Simulation Software
Benchmarking Techniques
Artificial Intelligence (AI)
Download MeeBoss