Accounting Expert (QuickBooks) — AI Model Evaluation
Contract · Remote · $120/hr
Role Overview
Huzzle is seeking experienced accountants to test how well frontier AI systems handle real accounting work. You will write realistic accounting prompts that are difficult enough to make a frontier model fail, and then write a rubric for each one that lets any reviewer grade the model's output consistently. A rubric is a list of objective, checkable criteria — a reviewer applies it to an answer and marks each criterion pass or fail.Training is provided.
Prior AI or data-annotation experience is preferred.
Key Responsibilities
- Write realistic accounting tasks of the kind a client, controller or business owner would send
- Month-end close, bank and credit card reconciliations, AR/AP ageing, accruals and adjusting entries, sales tax, job costing, and cleaning up a messy QuickBooks file
- Each request should call for a working deliverable, usually a spreadsheet, a journal entry schedule or a short memo.
- Run the request against a frontier AI model and examine the output with the scrutiny you would apply to work you were signing off. Where the model succeeds, escalate the difficulty until a genuine gap appears.
- Author a scoring rubric for each task with objective criteria another reviewer could apply without repeating your work, for example "the reconciled ending balance equals $X" or "the $X deposit is recorded as a customer deposit liability, not revenue."
Requirements
- 5+ years in accounting, bookkeeping or finance operations
- 3-4+ years of hands-on QuickBooks (Online or Desktop) in a real business setting
- Available 40 hours per week for the next 2 weeks
- Excellent written English and precision with numbers
Contract and Payment Terms You will be engaged as an independent contractor. Fully remote, completed on your own schedule within the engagement window. The engagement may be extended, shortened or concluded early depending on programme needs and performance.
About Huzzle
Huzzle partners with frontier AI labs to evaluate and improve frontier models using deep human expertise. Contributors work directly on the assessment of advanced AI systems in their own field, are paid competitively, and help define the standards by which those systems are measured.