AI Benchmark Engineer — Software Engineering

turing

Remote 5 Years Exp Posted 8d ago

Job Description

  • Build multi-agent benchmark tasks based on real-world open-source code changes (bug fixes, migrations, refactors)
  • Work with the Harbor evaluation framework to run and validate tasks inside Docker environments
  • Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints
  • Design and implement Python-based verification scripts to validate correctness of agent-generated code changes
  • Create decomposition strategies that split complex code changes across multiple independent sub-agents
  • Run, debug, and refine tasks within containerized environments to ensure reproducibility and determinism
  • Evaluate task performance signals and improve task quality, clarity, and difficulty

Requirements

  • 5+ years of experience in Python and JavaScript development
  • Experience with AI coding benchmarks (e.g., SWE-bench, Terminal-Bench)
  • Strong experience reading and navigating large open-source codebases (e.g., Django, Flask, FastAPI, Node.js, or similar)
  • Familiarity with Git workflows, including pull requests, diffs, cherry-picking, and working with specific commits
  • Comfortable working with Docker (writing Dockerfiles, building images, debugging container issues)
  • Experience writing test scripts (pytest, unittest, or custom assertion-based testing)
    • Ability to write clear, precise, and unambiguous technical specifications

Similar Openings for You