Role Overview
Huzzle is seeking finance professionals to test how well frontier AI systems handle real financial work, and to document where they fall short.
You will not be performing financial analysis.
You will be finding the point at which an AI system stops being reliable at it, and proving that point objectively.
Each piece of work you submit has three parts: a realistic financial request, documented evidence of where the model failed it, and a scoring rubric that lets any reviewer grade a future attempt consistently.
Full training and worked examples are provided.
Prior AI or data-annotation experience is preferred.
Key Responsibilities
- Test financial requests against a frontier AI model and examine its output with the scrutiny you would apply to work you were signing off.
- Document exactly where and how it failed, with the cell reference, figure or omission that demonstrates it.
- Precision matters more than volume.
- Author scoring rubrics with sourced, objective criteria that another reviewer could apply without repeating your research.
- Write the underlying requests: realistic financial tasks requiring live research against public sources, multi-step calculation, and an actual file deliverable — typically an Excel workbook, but also memos, PDFs, decks or CSVs.
- Escalate task difficulty where the model succeeds, until a genuine gap is exposed.
Ideal Qualifications
- 3+ years in financial analysis, FP&A, corporate finance, audit, tax, accounting or financial control.
- Strong Excel modelling ability — multi-tab workbooks, driver-based logic, scenario modelling.
- Not template completion.
- The instinct to verify a figure against its source rather than accept it.
- Excellent written English and unusual precision with detail.
- 15–25 hours per week available across the engagement.
- Part-qualified or qualified (ACA, ACCA, CIMA, CFA) is a plus, not a requirement.
- Demonstrated depth matters more than credentials.
Contract and Payment Terms
You will be engaged as an independent contractor.
Fully remote, completed on your own schedule within the engagement window.
The engagement may be extended, shortened or concluded early depending on programme needs and performance.
About Huzzle
Huzzle partners with frontier AI labs to evaluate and improve frontier models using deep human expertise.
Contributors work directly on the assessment of advanced AI systems in their own field, are paid competitively, and help define the standards by which those systems are measured.