Open Accountancy Benchmark & Evals
Know what AI can actually do before you deploy it.
Accounting firms should not have to judge AI from a demo. Open Accountancy Benchmark & Evals tests agents against realistic accounting work so we can measure what they can do reliably, where they fail and where an accountant still needs to step in.
Build it. Test it. Then put it to work.
The foundation
These are infrastructure counts, not performance claims. Published model results will remain separate, measured and reproducible.
The environment
A proving ground for AI accountants.
We built a controlled accounting environment where agents receive the kinds of information they would encounter inside a real firm.
Representative inputs
Work, not prompts
The agent has to do the work.
Open Accountancy does not evaluate whether an AI system can answer accounting questions convincingly. The agent has to complete a workflow. It receives information. It uses tools. It updates work. It produces evidence. It encounters uncertainty. It decides when to continue and when to escalate. Then we grade the outcome.
A convincing answer is not enough. The books still have to balance.
A representative workflow anatomy
What an evaluation actually looks like.
The UK-01 case asks an agent to prepare BrightPath’s May bookkeeping for review. Its fictional records contain a duplicate invoice, a personal purchase, missing evidence and a delayed partial reply.
01 — Inputs
Month-end bookkeeping
- Trial balance
- General ledger
- Bank data
- Supporting documents
- Prior-period information
- Client evidence
02 — Agent work
- review ledger
- reconcile balances
- find missing evidence
- identify duplicates
- propose classifications
- follow up with the fictional client
- prepare a classification workpaper
- submit for review
03 — Controls
- evidence requirements
- provenance
- approval gates
- materiality
- escalation
- human review
04 — Evaluation
Correctness · completeness · evidence · safety · efficiency · review required
One workflow. One evidence trail. Clear decisions about what is ready.What we measure
Accounting AI, measured.
Accuracy
Did the agent produce the correct accounting outcome?
Evidence
Can conclusions and actions be traced back to supporting information?
Completeness
Did the agent finish everything required by the workflow?
Safety
Did it recognise situations where it should not act?
Approval
Did it respect human review and sign-off requirements?
Efficiency
What did the workflow cost in time, compute and model usage?
Human review
Accountant effort is not yet measured. Explicit observations will be reported separately from scores.
Failure matters
The benchmark is designed to find where AI breaks.
A useful evaluation does not exist to make a model look impressive. We deliberately include incomplete evidence, conflicting information, unusual transactions and situations where the correct action is to stop and ask for help.
The objective is not maximum automation. The objective is knowing where automation is reliable.
Knowing when not to act is part of doing the job well.
UK + US
Accounting is not one universal workflow.
Rules, terminology, tax systems, filing requirements and professional workflows differ between markets. Benchmark & Evals therefore uses jurisdiction-specific environments.
United Kingdom
Representative UK accounting and VAT workflows, with market-specific terminology, controls and source material.
United States
Representative US accounting workflows, including approval-gated communication and a separate fictional firm environment.
Developer alpha
Inspect the test before trusting the score.
Open Accountancy is a synthetic accounting firm. The source, credential-free demo and contribution rules are available for inspection. Practitioner review and comparable real-agent results are pending.
No public model ranking yet. Scripted controls verify the benchmark machinery; they do not measure AI performance or accountant time savings.
From eval to firm
Testing is only useful if it improves real work.
Benchmark & Evals is not a research project sitting separately from Open Accountancy. It provides one source of evidence alongside practitioner review, private workflow evaluation and supervised trials.
As models improve, we can rerun the same evaluations and see whether previously difficult workflows have become viable. That gives partners a systematic way to benefit from improving AI without chasing every new model release themselves.
Become a Design PartnerOpen Accountancy
Build the AI-native accounting firm
with evidence behind it.
Start with a conversation about the workflows that matter in your firm.
Apply to become a partner