Long-term contract and Fully Remote opening for an AI Evaluation Software Engineer to build and scale tooling that measures the performance of AI-powered software development tools. This hands-on role involves developing evaluation harnesses, automating benchmark runs, and ensuring results are accurate, reproducible, and aligned with human judgment.
Responsibilities
- Build evaluation harnesses and automation for software development use cases, including converting merged pull requests into repeatable benchmark tasks.
- Develop versioned evaluation processes using pinned dependencies, containerized runs, and isolated worktrees.
- Validate and calibrate evaluation methods against human judgment.
- Support execution-based benchmarking across software quality, productivity, cost, and latency.
- Analyze repeated runs to identify variance, failure patterns, and cost per outcome.
- Partner with engineering and data teams to improve tooling and document methodology and results for technical and leadership audiences.
Required Skills
- Strong software engineering background with experience building automation, developer tooling, or test and validation systems.
- Understanding of LLMs, generative AI, data science, and machine learning fundamentals.
- Ability to create AI skills and design routing, orchestration, and prompt patterns, with supporting documentation and dependency management.
- Background in knowledge graphs, search, API design, and integrations.
Preferred Experience
- AI-powered coding tools and agentic applications such as Claude Code, Devin, or OpenCode.
- Software benchmarks and evaluation systems, especially execution-based grading against tests.
- LLM-as-judge or agent-as-judge approaches and validation against human raters.
- Build-system-aware test selection, including Bazel or mapping changed files to relevant tests.
- Reproducible test environments and versioned evaluation datasets.
- Communicating evaluation methodology and findings to engineering leadership.