AI Benchmark Engineer — Software Engineering
turing
Job Description
- Build multi-agent benchmark tasks based on real-world open-source code changes (bug fixes, migrations, refactors)
- Work with the Harbor evaluation framework to run and validate tasks inside Docker environments
- Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints
- Design and implement Python-based verification scripts to validate correctness of agent-generated code changes
- Create decomposition strategies that split complex code changes across multiple independent sub-agents
- Run, debug, and refine tasks within containerized environments to ensure reproducibility and determinism
- Evaluate task performance signals and improve task quality, clarity, and difficulty
Requirements
- 5+ years of experience in Python and JavaScript development
- Experience with AI coding benchmarks (e.g., SWE-bench, Terminal-Bench)
- Strong experience reading and navigating large open-source codebases (e.g., Django, Flask, FastAPI, Node.js, or similar)
- Familiarity with Git workflows, including pull requests, diffs, cherry-picking, and working with specific commits
- Comfortable working with Docker (writing Dockerfiles, building images, debugging container issues)
- Experience writing test scripts (pytest, unittest, or custom assertion-based testing)
- Ability to write clear, precise, and unambiguous technical specifications