All Jobs Vacancy

AI Evaluation Software Engineer

Posted 1 day ago by IDR, Inc.

Long-term contract and Fully Remote opening for an AI Evaluation Software Engineer to build and scale tooling that measures the performance of AI-powered software development tools. This hands-on role involves developing evaluation harnesses, automating benchmark runs, and ensuring results are accurate, reproducible, and aligned with human judgment.

Responsibilities

  • Build evaluation harnesses and automation for software development use cases, including converting merged pull requests into repeatable benchmark tasks.
  • Develop versioned evaluation processes using pinned dependencies, containerized runs, and isolated worktrees.
  • Validate and calibrate evaluation methods against human judgment.
  • Support execution-based benchmarking across software quality, productivity, cost, and latency.
  • Analyze repeated runs to identify variance, failure patterns, and cost per outcome.
  • Partner with engineering and data teams to improve tooling and document methodology and results for technical and leadership audiences.

Required Skills

  • Strong software engineering background with experience building automation, developer tooling, or test and validation systems.
  • Understanding of LLMs, generative AI, data science, and machine learning fundamentals.
  • Ability to create AI skills and design routing, orchestration, and prompt patterns, with supporting documentation and dependency management.
  • Background in knowledge graphs, search, API design, and integrations.

Preferred Experience

  • AI-powered coding tools and agentic applications such as Claude Code, Devin, or OpenCode.
  • Software benchmarks and evaluation systems, especially execution-based grading against tests.
  • LLM-as-judge or agent-as-judge approaches and validation against human raters.
  • Build-system-aware test selection, including Bazel or mapping changed files to relevant tests.
  • Reproducible test environments and versioned evaluation datasets.
  • Communicating evaluation methodology and findings to engineering leadership.
Rate:
Not specified
Location:
Remote
IR35 Status:
Outside
Remote Status:
Remote
Industry:
AI & Machine Learning
Seniority Level:
Not Specified

Take-Home Pay

Not Available

Visit calculators for additional details

Share job