10/01/2026

AI

Introducing Financial Audit Bench: Evaluating Audit Capabilities of AI Agents

Strict pass rate versus mean cost per run for the 11 frontier models tested in Financial Audit Bench. Source: Modus Labs.

As more audit firms are putting AI agents to work, the question they keep running into is simple: which parts of an audit can an agent do well? Until now, there was no shared way to answer it.

Vendors could point to hours saved, but no one could point to a consistent way to verify whether the work was accurate.

When testing 11 frontier AI models on financial audit procedures, the models would pass more than 85% of the individual checks. But none fully passed more than 69% of the procedures it attempted. And because audit work has to hold up every time, not just once, each procedure was run eight times. When the top model was required to pass all eight runs, its success rate fell from 69.17% to 44.4%. On accounts receivable, the average pass rate was only 4%.

These results come from Financial Audit Bench (FAB), an open-source benchmark released by Modus Labs, the research team at the AI-native accounting firm Modus. FAB tests whether AI agents can complete the multi-step procedures of a financial statement audit. The results are fully public, including a live leaderboard.

We’ve seen benchmarks set the standard for AI in fields like accounting and finance. Audit didn’t have one until now with FAB.

Why audit merits needs its own benchmark

Many accounting firms are short on staff, and many turn to agents, whether or not the tools are ready. They need to know which parts agents can handle, and with what degree of correctness.

Audit quality is hard to guarantee even when humans do all the work. It’s estimated that 40% of audits contain significant errors. Auditors test samples, so much of a company’s activity never gets a direct look, which makes quality hard to measure.

“While time savings are straightforward to quantify, audit quality has historically been an area with less precise measures,” says Pranav Pillai, Chief Technology Officer at Modus.

The benchmark has 45 workpapers across eight audit areas, modeled on the fieldwork a staff-level auditor does. “At firms today, senior auditors spot-check key details in a workpaper rather than re-performing audit procedures. We mirror this process in FAB,” says Pranav. To protect client records, Modus generated six synthetic company audits from aggregate statistics on past audits.

To develop FAB, Modus Labs paired PhDs and applied AI engineers with veterans of public accounting, including advisor Jim Burton, former Chief Auditor at Grant Thornton.

Auditors put more than 1,100 hours into developing and reviewing the benchmark. Built with deep expertise, FAB reflects what a reviewer would look for.

What Financial Audit Bench found in the results

Agents passed journal-entry completeness checks 94.5% of the time. However, revenue testing came in at 8%, and accounts receivable at 4%. The weakest areas are the ones estimated to take auditors the most time.

Modus found that agents often stopped looking too early and trusted the wrong evidence. One agent tested receivables against the client’s own cash-receipts report when it should have used the bank statement.

The results only measure the tasks specifically scoped in the benchmark. Whether an agent can run an audit alone is a separate question. “We therefore recommend deploying agents as co-pilots rather than autonomous preparers,” says Pranav.

In the weeks Modus spent building the benchmark, the leading strict-pass rate rose from below 50% to 69.2%. Firms should build systems that can test and adopt stronger models as they become available.

What’s next for benchmarks with ever-changing models

Modus was founded to help deliver the highest-quality audit in a fraction of the time. FAB asks which audit tasks are well-suited for agents and which models suit those tasks. It also maps the mistakes a reviewer should expect to catch.

Correct answers are only one part of quality. Modus’s auditors found that agents often wrote long descriptions of their process but didn’t emphasize the key evidence or conclusions. Future evaluations should also measure whether another auditor can review and follow the work.

Modus has released the paper, code, dataset, and live leaderboard so anyone can rerun the benchmark and build on it. We believe open benchmarks are how an industry decides what it can trust AI to do.

Read Modus’s full announcement of Financial Audit Bench here.

The content here does not constitute an offer to sell or a solicitation of an offer to buy any securities or investment advisory services.

The views expressed are those of the authors and do not necessarily represent the views or opinions of Lightspeed. Other market participants could take different views. Unless otherwise indicated, the inclusion of any third-party firm and/or company names, brands and/or logos are for representational purposes and does not imply any affiliation with these firms or companies and also does not imply their endorsement of the views expressed by the authors.

Certain information contained herein is based on information from various sources prepared by third parties. While such sources are believed by Lightspeed to be reliable, neither Lightspeed nor its affiliates assume any responsibility for the accuracy or completeness of such information, and such information has not been independently verified by Lightspeed. For more details please see https://www.lsvp.com/legal.

Lightspeed Possibility grows the deeper you go. Serving bold builders of the future.