FAB: a benchmark for real financial work
We built a synthetic company, its books and a 160-file data room, then gave six models the work a deal team actually does. The highest pass rate was 62% across repeated task attempts.
Can AI do financial analysis we can trust?
AI agents can read financial documents, query accounting data, and write analysis. To be useful to financial analysts, they need to work through the evidence behind the numbers: reconcile conflicting sources, test management’s explanations, and produce conclusions that people can inspect, verify, and challenge.
We introduce FAB - Finance Agents Benchmark, an evaluation of how well AI agents answer financial due-diligence questions using a company’s books and supporting documents. The benchmark includes:
- A realistic company and data room. Fifty tasks use a synthetic industrial distributor’s financial records and 160 files, spanning SAP extracts, spreadsheets, contracts, presentations, and emails.
- Ground truth built into the books. We generate the company’s transactions before its documents. Nineteen planted issues flow through the accounting records and supporting evidence, so expected answers can be traced to their source.
- Complete answers and consistent performance. Each answer must pass every grading criterion, covering figures, reconciliations, conclusions, and gaps in the evidence. We evaluate six models across three trials per task, measuring both how often they succeed and whether they succeed consistently.
The code and run instructions are on GitHub, and the data room and tasks are on Hugging Face.
| Spec | Value |
|---|---|
| Company / industry | Meridian Industrial Supply LLC (synthetic) · Industrial distribution, United States |
| Periods | FY2024 and FY2025, evidence to 15 February 2026 |
| Ledger | ~22,000 journal lines |
| Data room | 160 files (SAP CSV, XLSX, PDF, DOCX, PPTX, EML) |
| Planted facts | 19 |
| Tasks / Criteria | 50 (15 easy, 20 medium, 15 hard) |
| Grading | 231 pass/fail criteria |
Most questions require evidence from several files. A revenue number might be in SAP, the reason it moved might be in a contract, and whether management's explanation holds up might depend on comparing both. That is the kind of work we wanted to test.
What a data room contains
When a company is being sold, the buyer's deal team gets access to a data room: the company's financial records and the documents behind them. A financial due-diligence analyst works through that room to answer questions like: what actually drove revenue growth? Are the EBITDA adjustments supportable? How much of the year-end cash is really available?
We built the benchmark around a synthetic company, Meridian Industrial Supply LLC, a US industrial distributor. Its data room has the same broad layers you would expect in a real process:
- The books. Raw SAP extracts of every posting, trial balances, monthly management accounts, receivables and payables ageing, bank statements.
- The evidence behind the books. Customer and supplier contracts, invoices, credit notes, the fixed asset register, the loan agreement, payroll summaries.
- What management says. The management presentation, a trading update, board minutes, the proposed EBITDA adjustments and emails with the deal team.
The job is not to trust one layer. It is to check them against each other.
A concrete example: freight booked in the wrong year
One planted issue is a year-end cut-off error. Two December 2025 freight invoices - $260,000 from Midwest Freight and $160,000 from Lakefront Logistics - are posted in January 2026 with no December accrual. That leaves FY2025 costs understated by $420,000.
Nothing in the room says "cut-off error". The agent has to notice that the costs were recorded in January, the invoices say the services happened in December, and there is no matching December accrual in payables.
We also put those invoices among normal freight invoices, including a correctly accrued December invoice from the same supplier. So pattern-matching on "December invoice = problem" is not enough.
That same error matters elsewhere. It reduces corrected EBITDA and, together with the other planted corrections, causes the company to breach its year-end loan covenant. Because the error is generated once in the underlying books, every document and every answer that depends on it stays consistent.
Building the company
The important design choice is that we do not write the documents first. We model the company’s transactions first, generate its books, and then render the documents from those books.
That gives us one source of truth instead of a pile of hand-authored files that can drift apart.
1. Start from a real company. We use a real company's financial profile as a baseline, then scale revenue, costs and working capital to fit Meridian. We build synthetic transactions and events around that baseline, then generate the ledger and SAP records.
2. An ordinary business. One specification file describes how Meridian normally runs: 6 customer accounts, 16 suppliers, 22 products, around 250 employees in 5 departments, monthly purchases, payroll, rent, a term loan and its covenant. Most of the ledger is intentionally boring.
3. Planted facts. On top of that ordinary activity, we add 19 specific issues we want the analyst to be able to find. Some are simple, such as revenue growth. Some are traps, such as a gross-margin improvement caused by a one-off supplier rebate while management credits "pricing efficiencies". Some are deliberately unresolvable: two customers share an address, but there are no ownership records in the room to prove they are related.
We also include correctly handled items alongside the problems. An agent that assumes management is always wrong should fail.
4. Generation. A generator turns the specification into a full double-entry ledger. Every posting comes from a modelled transaction, the books balance every month, and the same specification produces the same books every time.
5. The answer key. From the finished ledger, the generator computes every figure the tasks depend on - more than 1,000 values - and records which underlying entries produced each figure. The answer key is private and never enters the data room.
6. The data room. Finally, the 160 documents are rendered from the ledger and specification: SAP tables as CSV, schedules as spreadsheets, contracts and letters as PDF or Word, the management presentation as PowerPoint, and correspondence as email. If a document is supposed to be missing, it is simply not there.
Environment and tasks
All 50 tasks run against the same 160-file data room in an isolated sandbox with no network access. The agent receives the data room and a deal-team question. The answer key is withheld. The agent can read every format in the room and run Python or DuckDB over the SAP extracts. Each run starts from a fresh workspace and produces one answer file.
The tasks are written like requests from a deal team:
- Easy: one or two documents and a lookup or total. "What was FY2025 net revenue, and do SAP and the management accounts agree?"
- Medium: combine several documents into a reconciliation or bridge. "Are any FY2025 costs recorded in FY2026? Quantify them."
- Hard: judgement across conflicting evidence. "Assess each EBITDA adjustment and give a supported adjusted EBITDA."
Ground truth comes from the generated answer key, not from our reading of the documents after the fact. Each task is broken into individual pass/fail checks - a figure, a conclusion, a missing document the agent should request - and an LLM judge grades each criterion separately.
Results
We ran six models on all 50 tasks three times each, with gpt-6-luna at maximum reasoning as the judge. That produced 900 completed and scored answers. The explorer below shows complete answers, criteria passed, and how consistently each model gets a task right across repeated runs.
Complete answers / 150 trials. Every criterion must pass.
DeepSeek passes 90 of 150 answers (60.0%) and Sol passes 88 (58.7%). The overall scores are close, but performance differs by task difficulty.
DeepSeek is much stronger on the easy set, completing 42 of 45 easy answers versus Sol's 35. On hard tasks, that flips: Sol completes 21 of 45, while DeepSeek completes 13.
That split is more interesting to us than saying the two models are simply "tied."
Kimi K3 passes 93 of 150 answers (62.0%), with 83.1% of criteria passed. It passes 37 of 50 tasks at least once and 25 on all three trials. MiniMax M2.5 passes 37 of 150 answers (24.7%), with 52.8% of criteria passed. It passes 15 tasks at least once and 9 on all three trials.
All-three reliability (pass³): Kimi scores 50% (25/50 tasks). This is the share of tasks that pass every criterion on all three trials, distinct from its 62.0% run success rate. Sol scores 48%, DeepSeek 46%, Luna 38%, GLM 36% and MiniMax 18%. The metric is conventionally written passᵏ; here, k = 3.
Kimi scores 95.6% on easy tasks, 63.3% on medium and 26.7% on hard. MiniMax scores 73.3%, 6.7% and 0%, respectively.
The system prompt changed between the historical and Bedrock cohorts, so the comparison is not controlled for model differences alone. MiniMax’s scores include the three recovered context-limit failures; the original attempts remain recorded in the source results.
A few other things stood out:
- Across models, easy answers pass 73–96% of the time, medium 7–63%, and hard 0–47%.
- Models get a lot of individual pieces right without finishing the whole job: 53–83% of criteria pass, but only 25–62% of full answers do.
- Reliability is still weak. Models pass 30–76% of tasks at least once, but only 18–50% on all three trials.
DeepSeek
28.9%13/45 complete answers66.3% of criteria passSol
46.7%21/45 complete answers72.9% of criteria passLuna
28.9%13/45 complete answers66.3% of criteria passGLM
22.2%10/45 complete answers58.5% of criteria passKimi
26.7%12/45 complete answers68.2% of criteria passMiniMax
0.0%0/45 complete answers26.4% of criteria passExplore the tasks
Eight tasks were passed on every run, including reported revenue, reported EBITDA and the customer cohort analysis. Those tasks are useful checks, but they no longer separate the models very much.
Seven tasks were never passed by any model: monthly working capital, closing debt-like items, supported adjusted EBITDA, the normalised working-capital peg, run-rate EBITDA, the downside from customer concentration and the enterprise-to-equity bridge.
Many failed answers still identify several of the relevant facts. They may miss one part of the judgement or reconciliation. That is exactly the gap we care about: being partly right is not the same as producing work a deal team can rely on.
A failed answer: leaving a $720,000 write-down out of EBITDA
Task 038 asks the agent to assess each EBITDA adjustment and produce a supported total. In its second trial, DeepSeek V4.1 Flash passed 10 of the 13 checks. It correctly accepted the $900,000 system-implementation and $650,000 settlement add-backs, rejected recurring severance, and deducted the $420,000 freight correction described above.
It also found the inventory problem. Meridian held 6,000 obsolete seal packs at a recorded cost of $900,000. The stock committee minutes showed no customer demand since June 2023, and a salvage quote valued the packs at $180,000. DeepSeek calculated the resulting $720,000 write-down correctly, then explicitly left it out of the EBITDA bridge because it treated the write-down only as a balance-sheet and working-capital adjustment. The benchmark required that omitted FY2025 expense to reduce EBITDA as well.
The bonus adjustment went wrong in the other direction. DeepSeek treated the $1.2 million retention pool as a separate, wholly unrecorded obligation and deducted it in full. The benchmark's company specification and answer key treat the $600,000 already accrued in payroll as part of that pool, leaving a further $600,000 expense to recognise.
The two errors partly cancelled out: omitting the inventory expense overstated EBITDA by $720,000, while the excess bonus deduction understated it by $600,000. DeepSeek reported $21.576 million against the expected $21.456 million, both including the conditional $480,000 rent saving. It failed the inventory, bonus and final-total checks. The $120,000 gap in the final figure hid two larger errors that a reviewer would need to correct before relying on the analysis.
50 of 50 tasks · Cells show complete answers out of three.
| Task | DeepSeek | Sol | Luna | GLM | Kimi | MiniMax |
|---|---|---|---|---|---|---|
| 001 / easyFY2025 net revenue reconciliation | ||||||
| 002 / easyLargest customer accounts | ||||||
| 003 / easyOverdue receivables | ||||||
| 004 / easyManagement proposed EBITDA adjustments | ||||||
| 005 / easyLargest supplier and contract | ||||||
| 006 / easyMonthly revenue outlier | ||||||
| 007 / easyClosing cash and bank reconciliation | ||||||
| 008 / easyHeadcount and payroll | ||||||
| 009 / easyReported EBITDA | ||||||
| 010 / easyReported gross margins | ||||||
| 011 / easyInventory balances and days | ||||||
| 012 / easyYear-end DSO | ||||||
| 013 / easyCustomer credits | ||||||
| 014 / easyOperating expense categories | ||||||
| 015 / easyActive customer counts | ||||||
| 016 / mediumRevenue growth by customer cohort | ||||||
| 017 / mediumCustomer concentration and control | ||||||
| 018 / mediumManagement accounts and trial balance | ||||||
| 019 / mediumFreight cutoff correction | ||||||
| 020 / mediumMonthly operating working capital | ||||||
| 021 / mediumOverdue customer balance and deal treatment | ||||||
| 022 / mediumSlow stock adjustment | ||||||
| 023 / mediumClosing debt-like items | ||||||
| 024 / mediumRelated-party transactions | ||||||
| 025 / mediumGross-margin bridge | ||||||
| 026 / mediumMonthly revenue pattern and unusual sale | ||||||
| 027 / mediumPayment days and year-end stretch | ||||||
| 028 / mediumSeverance and recurrence | ||||||
| 029 / mediumActual versus budget | ||||||
| 030 / mediumSubsequent credits and cutoff | ||||||
| 031 / mediumKestrel payment terms | ||||||
| 032 / mediumCapex versus depreciation | ||||||
| 033 / mediumYear-end covenant compliance | ||||||
| 034 / mediumBonus correction and net debt | ||||||
| 035 / mediumWarehouse rent normalisation | ||||||
| 036 / hardUnusual December sale | ||||||
| 037 / hardManagement growth and run-rate claims | ||||||
| 038 / hardSupported adjusted EBITDA | ||||||
| 039 / hardManagement margin explanation | ||||||
| 040 / hardEarnings to operating and free cash flow | ||||||
| 041 / hardNormalised working-capital peg | ||||||
| 042 / hardDefensible run-rate EBITDA | ||||||
| 043 / hardKestrel dependence and downside | ||||||
| 044 / hardSupplier economics and margin sustainability | ||||||
| 045 / hardOwner-related normalisations | ||||||
| 046 / hardYear-end presentation and cash timing | ||||||
| 047 / hardSupported covenant headroom | ||||||
| 048 / hardMissing evidence requests | ||||||
| 049 / hardEnterprise to equity bridge | ||||||
| 050 / hardDeal-team red-flag memo |
TASK 019 · medium · GPT-6 Luna
Freight cutoff correction
What adjustment to FY2025 EBITDA do year-end cost cutoff issues require?
| Required check | Trial 1 | Trial 2 | Trial 3 |
|---|---|---|---|
| Midwest freight | Pass | Pass | Pass |
| Lakefront freight | Pass | Pass | Pass |
| EBITDA correction | Pass | Pass | Pass |
| Scope and precision | Pass | Pass | Pass |
| Complete answer | Pass | Pass | Pass |
Saved judge verdicts, not an independent human review. Failure of one check fails the task.
Resource use
Estimated cost without caching · USD per completed task run
The estimated cost per completed task run is $0.064 for GPT-6 Luna, $0.138 for GLM 5.3 Flash, $0.881 for DeepSeek V4.1 Flash and $0.945 for GPT-6 Sol, excluding judging and provider caching. Among the original four models, DeepSeek is the least token-efficient: it uses roughly 3 to 7 times as many tokens as the other models and about twice as many turns as Sol.
Kimi averages 2.63 million tokens and 25.9 turns per completed run, close to DeepSeek’s 2.87 million tokens and 27.5 turns. MiniMax averages 0.59 million tokens and 18.8 turns. Using AWS Bedrock Standard rates for their recorded inference routes, estimated cost per completed run is $8.982 for Kimi and $0.183 for MiniMax.
AWS lists Kimi K3’s US cross-region rates at $3.30 per million input tokens and $16.50 per million output tokens. MiniMax M2.5’s Oregon rates are $0.30 and $1.20, respectively. These estimates treat input as uncached, matching the original cost comparison; they do not represent actual bills.
Higher and further left is better. The dashed Pareto frontier connects models for which no other model is both at least as accurate and at least as cheap, with one strict improvement.
DeepSeek V4.1 Flash: $0.881 per completed run · 60.0% pass rate · 2.849M input tokens + 22.1K output tokens · 27.5 agent turns on average.
| Model | USD / run | Task pass rate | USD / passing answer |
|---|---|---|---|
| GPT-6 Luna | $0.064 | 50.7% | $0.126 |
| GLM 5.3 Flash | $0.138 | 47.3% | $0.292 |
| MiniMax M2.5 | $0.183 | 24.7% | $0.744 |
| DeepSeek V4.1 Flash | $0.881 | 60.0% | $1.469 |
| GPT-6 Sol | $0.945 | 58.7% | $1.610 |
| Kimi K3 | $8.982 | 62.0% | $14.487 |
AWS rates and calculation
USD per million tokens, verified 2026-09-29. These rates match the inference routes in the saved runs.
| Model / route | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| Kimi K3US cross-region · Standard | $3.30 | $16.50 | $0.33 | $4.125 |
| MiniMax M2.5Oregon · Standard on-demand | $0.30 | $1.20 | Not listed | Not listed |
For Kimi and MiniMax: mean cost = (completed-run input tokens × input rate + completed-run output tokens × output rate) / 1,000,000 / 150. Each completed trial counts, whether its answer passes or fails. Superseded attempts, execution failures, judging and infrastructure are excluded.
All input is priced as uncached for consistency with the original comparison. Kimi supports automatic caching, but the saved usage does not split cache reads and writes, so these are estimates, not invoices. The listed cache rates are not applied; Kimi’s cache-write rate is for a 30-minute retention period.
The original four models retain the report’s OpenAI Standard and Fireworks serverless pricing snapshots from 24–25 September 2026, including GPT-6 Sol’s long-context pricing adjustment.
Under these assumptions, Luna, DeepSeek and Kimi form the observed cost-efficiency frontier. Luna has the lowest estimated cost per run, DeepSeek adds a higher pass rate, and Kimi’s two-point increase over DeepSeek comes at roughly ten times the estimated cost. Sol, GLM and MiniMax each have another model that costs less and passes more often in these results. The different system prompts remain a limitation of this comparison.
Limitations
These are first results, not a final ranking of models for financial work.
All tasks share one company, so they are not independent samples. A failed answer means the response did not satisfy every criterion; it does not tell us how much work a human reviewer would need to fix it. And every criterion is graded by one judge from the same family as one of the evaluated models, calibrated on authored reference and flawed answers rather than practitioner review.
Tasks 041 and 049 are included using their saved regrades. We had previously excluded them because their questions rely on assumptions that could be stated more clearly; including them does not resolve that concern.
We plan to address these limitations before making broader claims.
Conclusion
The main thing we learned is that deal-side finance is a useful stress test for agents because the work is not just arithmetic. The evidence is spread across formats, the numbers have to reconcile, and the important finding is often something nobody states directly.
The benchmark only works because the company comes before the documents. Every planted fact has one source of truth, every downstream file is generated from it, and every expected answer can be traced back to the underlying records.
That also gives us a path forward: build more companies with different stories, not just more questions about Meridian. Next we are expanding the benchmark and adding practitioner review to the grading process.