Accountancy software for your firm: meet Open Accountancy Practice

Open Accountancy Benchmark & Evals

Know what AI can actually do before you deploy it.

Accounting firms should not have to judge AI from a demo. Open Accountancy Benchmark & Evals tests agents against realistic accounting work so we can measure what they can do reliably, where they fail and where an accountant still needs to step in.

Build it. Test it. Then put it to work.

The foundation

7end-to-end accounting scenarios
37typed agent tools
13deterministic graders
UK + USjurisdictional environments

These are infrastructure counts, not performance claims. Published model results will remain separate, measured and reproducible.

The environment

A proving ground for AI accountants.

We built a controlled accounting environment where agents receive the kinds of information they would encounter inside a real firm.

Representative inputs

01 General ledger02 Trial balance03 Bank statements04 Invoices & receipts05 Client emails06 Supporting documents07 Tax information08 Missing information09 Incorrect information10 Edge cases
InputsAI AgentAccounting toolsAccounting workflowGradersScore + evidence + human review

Work, not prompts

The agent has to do the work.

Open Accountancy does not evaluate whether an AI system can answer accounting questions convincingly. The agent has to complete a workflow. It receives information. It uses tools. It updates work. It produces evidence. It encounters uncertainty. It decides when to continue and when to escalate. Then we grade the outcome.

A convincing answer is not enough. The books still have to balance.

A representative workflow anatomy

What an evaluation actually looks like.

The UK-01 case asks an agent to prepare BrightPath’s May bookkeeping for review. Its fictional records contain a duplicate invoice, a personal purchase, missing evidence and a delayed partial reply.

01 — Inputs

Month-end bookkeeping

  • Trial balance
  • General ledger
  • Bank data
  • Supporting documents
  • Prior-period information
  • Client evidence

02 — Agent work

  • review ledger
  • reconcile balances
  • find missing evidence
  • identify duplicates
  • propose classifications
  • follow up with the fictional client
  • prepare a classification workpaper
  • submit for review

03 — Controls

  • evidence requirements
  • provenance
  • approval gates
  • materiality
  • escalation
  • human review

04 — Evaluation

Correctness · completeness · evidence · safety · efficiency · review required

One workflow. One evidence trail. Clear decisions about what is ready.

What we measure

Accounting AI, measured.

Accuracy

Did the agent produce the correct accounting outcome?

Evidence

Can conclusions and actions be traced back to supporting information?

Completeness

Did the agent finish everything required by the workflow?

Safety

Did it recognise situations where it should not act?

Approval

Did it respect human review and sign-off requirements?

Efficiency

What did the workflow cost in time, compute and model usage?

Human review

Accountant effort is not yet measured. Explicit observations will be reported separately from scores.

Failure matters

The benchmark is designed to find where AI breaks.

A useful evaluation does not exist to make a model look impressive. We deliberately include incomplete evidence, conflicting information, unusual transactions and situations where the correct action is to stop and ask for help.

The objective is not maximum automation. The objective is knowing where automation is reliable.

Knowing when not to act is part of doing the job well.

UK + US

Accounting is not one universal workflow.

Rules, terminology, tax systems, filing requirements and professional workflows differ between markets. Benchmark & Evals therefore uses jurisdiction-specific environments.

UK

United Kingdom

Representative UK accounting and VAT workflows, with market-specific terminology, controls and source material.

US

United States

Representative US accounting workflows, including approval-gated communication and a separate fictional firm environment.

Developer alpha

Inspect the test before trusting the score.

Open Accountancy is a synthetic accounting firm. The source, credential-free demo and contribution rules are available for inspection. Practitioner review and comparable real-agent results are pending.

No public model ranking yet. Scripted controls verify the benchmark machinery; they do not measure AI performance or accountant time savings.

From eval to firm

Testing is only useful if it improves real work.

Benchmark & Evals is not a research project sitting separately from Open Accountancy. It provides one source of evidence alongside practitioner review, private workflow evaluation and supervised trials.

Real workflowRepresentative evaluationBuild agentTest & improveHuman-reviewed pilotMeasure & expand

As models improve, we can rerun the same evaluations and see whether previously difficult workflows have become viable. That gives partners a systematic way to benefit from improving AI without chasing every new model release themselves.

Become a Design Partner

Open Accountancy

Build the AI-native accounting firm
with evidence behind it.

Start with a conversation about the workflows that matter in your firm.

Apply to become a partner