FAB: a benchmark for real financial work

    We built a synthetic company, its books and a 160-file data room, then gave six models the work a deal team actually does. The highest pass rate was 62% across repeated task attempts.

    Can AI do financial analysis we can trust?

    AI agents can read financial documents, query accounting data, and write analysis. To be useful to financial analysts, they need to work through the evidence behind the numbers: reconcile conflicting sources, test management’s explanations, and produce conclusions that people can inspect, verify, and challenge.

    We introduce FAB - Finance Agents Benchmark, an evaluation of how well AI agents answer financial due-diligence questions using a company’s books and supporting documents. The benchmark includes:

    • A realistic company and data room. Fifty tasks use a synthetic industrial distributor’s financial records and 160 files, spanning SAP extracts, spreadsheets, contracts, presentations, and emails.
    • Ground truth built into the books. We generate the company’s transactions before its documents. Nineteen planted issues flow through the accounting records and supporting evidence, so expected answers can be traced to their source.
    • Complete answers and consistent performance. Each answer must pass every grading criterion, covering figures, reconciliations, conclusions, and gaps in the evidence. We evaluate six models across three trials per task, measuring both how often they succeed and whether they succeed consistently.

    The code and run instructions are on GitHub, and the data room and tasks are on Hugging Face.

    FAB benchmark specification
    SpecValue
    Company / industryMeridian Industrial Supply LLC (synthetic) · Industrial distribution, United States
    PeriodsFY2024 and FY2025, evidence to 15 February 2026
    Ledger~22,000 journal lines
    Data room160 files (SAP CSV, XLSX, PDF, DOCX, PPTX, EML)
    Planted facts19
    Tasks / Criteria50 (15 easy, 20 medium, 15 hard)
    Grading231 pass/fail criteria

    Most questions require evidence from several files. A revenue number might be in SAP, the reason it moved might be in a contract, and whether management's explanation holds up might depend on comparing both. That is the kind of work we wanted to test.

    What a data room contains

    When a company is being sold, the buyer's deal team gets access to a data room: the company's financial records and the documents behind them. A financial due-diligence analyst works through that room to answer questions like: what actually drove revenue growth? Are the EBITDA adjustments supportable? How much of the year-end cash is really available?

    We built the benchmark around a synthetic company, Meridian Industrial Supply LLC, a US industrial distributor. Its data room has the same broad layers you would expect in a real process:

    • The books. Raw SAP extracts of every posting, trial balances, monthly management accounts, receivables and payables ageing, bank statements.
    • The evidence behind the books. Customer and supplier contracts, invoices, credit notes, the fixed asset register, the loan agreement, payroll summaries.
    • What management says. The management presentation, a trading update, board minutes, the proposed EBITDA adjustments and emails with the deal team.

    The job is not to trust one layer. It is to check them against each other.

    A concrete example: freight booked in the wrong year

    One planted issue is a year-end cut-off error. Two December 2025 freight invoices - $260,000 from Midwest Freight and $160,000 from Lakefront Logistics - are posted in January 2026 with no December accrual. That leaves FY2025 costs understated by $420,000.

    Nothing in the room says "cut-off error". The agent has to notice that the costs were recorded in January, the invoices say the services happened in December, and there is no matching December accrual in payables.

    We also put those invoices among normal freight invoices, including a correctly accrued December invoice from the same supplier. So pattern-matching on "December invoice = problem" is not enough.

    That same error matters elsewhere. It reduces corrected EBITDA and, together with the other planted corrections, causes the company to breach its year-end loan covenant. Because the error is generated once in the underlying books, every document and every answer that depends on it stays consistent.

    Building the company

    The important design choice is that we do not write the documents first. We model the company’s transactions first, generate its books, and then render the documents from those books.

    That gives us one source of truth instead of a pile of hand-authored files that can drift apart.

    1. Start from a real company. We use a real company's financial profile as a baseline, then scale revenue, costs and working capital to fit Meridian. We build synthetic transactions and events around that baseline, then generate the ledger and SAP records.

    2. An ordinary business. One specification file describes how Meridian normally runs: 6 customer accounts, 16 suppliers, 22 products, around 250 employees in 5 departments, monthly purchases, payroll, rent, a term loan and its covenant. Most of the ledger is intentionally boring.

    3. Planted facts. On top of that ordinary activity, we add 19 specific issues we want the analyst to be able to find. Some are simple, such as revenue growth. Some are traps, such as a gross-margin improvement caused by a one-off supplier rebate while management credits "pricing efficiencies". Some are deliberately unresolvable: two customers share an address, but there are no ownership records in the room to prove they are related.

    We also include correctly handled items alongside the problems. An agent that assumes management is always wrong should fail.

    4. Generation. A generator turns the specification into a full double-entry ledger. Every posting comes from a modelled transaction, the books balance every month, and the same specification produces the same books every time.

    5. The answer key. From the finished ledger, the generator computes every figure the tasks depend on - more than 1,000 values - and records which underlying entries produced each figure. The answer key is private and never enters the data room.

    6. The data room. Finally, the 160 documents are rendered from the ledger and specification: SAP tables as CSV, schedules as spreadsheets, contracts and letters as PDF or Word, the management presentation as PowerPoint, and correspondence as email. If a document is supposed to be missing, it is simply not there.

    Environment and tasks

    All 50 tasks run against the same 160-file data room in an isolated sandbox with no network access. The agent receives the data room and a deal-team question. The answer key is withheld. The agent can read every format in the room and run Python or DuckDB over the SAP extracts. Each run starts from a fresh workspace and produces one answer file.

    The tasks are written like requests from a deal team:

    • Easy: one or two documents and a lookup or total. "What was FY2025 net revenue, and do SAP and the management accounts agree?"
    • Medium: combine several documents into a reconciliation or bridge. "Are any FY2025 costs recorded in FY2026? Quantify them."
    • Hard: judgement across conflicting evidence. "Assess each EBITDA adjustment and give a supported adjusted EBITDA."

    Ground truth comes from the generated answer key, not from our reading of the documents after the fact. Each task is broken into individual pass/fail checks - a figure, a conclusion, a missing document the agent should request - and an LLM judge grades each criterion separately.

    Results

    We ran six models on all 50 tasks three times each, with gpt-6-luna at maximum reasoning as the judge. That produced 900 completed and scored answers. The explorer below shows complete answers, criteria passed, and how consistently each model gets a task right across repeated runs.

    01 / MODEL COMPARISON6 models · 900 trials

    Complete answers / 150 trials. Every criterion must pass.

    Kimi K3
    62.0%93/150
    DeepSeek V4.1 Flash
    60.0%90/150
    GPT-6 Sol
    58.7%88/150
    GPT-6 Luna
    50.7%76/150
    GLM 5.3 Flash
    47.3%71/150
    MiniMax M2.5
    24.7%37/150
    Same room and calibrated rubric. Three fresh trials per task.

    DeepSeek passes 90 of 150 answers (60.0%) and Sol passes 88 (58.7%). The overall scores are close, but performance differs by task difficulty.

    DeepSeek is much stronger on the easy set, completing 42 of 45 easy answers versus Sol's 35. On hard tasks, that flips: Sol completes 21 of 45, while DeepSeek completes 13.

    That split is more interesting to us than saying the two models are simply "tied."

    Kimi K3 passes 93 of 150 answers (62.0%), with 83.1% of criteria passed. It passes 37 of 50 tasks at least once and 25 on all three trials. MiniMax M2.5 passes 37 of 150 answers (24.7%), with 52.8% of criteria passed. It passes 15 tasks at least once and 9 on all three trials.

    All-three reliability (pass³): Kimi scores 50% (25/50 tasks). This is the share of tasks that pass every criterion on all three trials, distinct from its 62.0% run success rate. Sol scores 48%, DeepSeek 46%, Luna 38%, GLM 36% and MiniMax 18%. The metric is conventionally written passᵏ; here, k = 3.

    Kimi scores 95.6% on easy tasks, 63.3% on medium and 26.7% on hard. MiniMax scores 73.3%, 6.7% and 0%, respectively.

    The system prompt changed between the historical and Bedrock cohorts, so the comparison is not controlled for model differences alone. MiniMax’s scores include the three recovered context-limit failures; the original attempts remain recorded in the source results.

    A few other things stood out:

    • Across models, easy answers pass 73–96% of the time, medium 7–63%, and hard 0–47%.
    • Models get a lot of individual pieces right without finishing the whole job: 53–83% of criteria pass, but only 25–62% of full answers do.
    • Reliability is still weak. Models pass 30–76% of tasks at least once, but only 18–50% on all three trials.
    02 / DIFFICULTY45 trials per model

    DeepSeek

    28.9%13/45 complete answers66.3% of criteria pass

    Sol

    46.7%21/45 complete answers72.9% of criteria pass

    Luna

    28.9%13/45 complete answers66.3% of criteria pass

    GLM

    22.2%10/45 complete answers58.5% of criteria pass

    Kimi

    26.7%12/45 complete answers68.2% of criteria pass

    MiniMax

    0.0%0/45 complete answers26.4% of criteria pass
    Labels were assigned before the runs. Harder tasks also contain more criteria; all-pass rates reflect both demands.

    Explore the tasks

    Eight tasks were passed on every run, including reported revenue, reported EBITDA and the customer cohort analysis. Those tasks are useful checks, but they no longer separate the models very much.

    Seven tasks were never passed by any model: monthly working capital, closing debt-like items, supported adjusted EBITDA, the normalised working-capital peg, run-rate EBITDA, the downside from customer concentration and the enterprise-to-equity bridge.

    Many failed answers still identify several of the relevant facts. They may miss one part of the judgement or reconciliation. That is exactly the gap we care about: being partly right is not the same as producing work a deal team can rely on.

    A failed answer: leaving a $720,000 write-down out of EBITDA

    Task 038 asks the agent to assess each EBITDA adjustment and produce a supported total. In its second trial, DeepSeek V4.1 Flash passed 10 of the 13 checks. It correctly accepted the $900,000 system-implementation and $650,000 settlement add-backs, rejected recurring severance, and deducted the $420,000 freight correction described above.

    It also found the inventory problem. Meridian held 6,000 obsolete seal packs at a recorded cost of $900,000. The stock committee minutes showed no customer demand since June 2023, and a salvage quote valued the packs at $180,000. DeepSeek calculated the resulting $720,000 write-down correctly, then explicitly left it out of the EBITDA bridge because it treated the write-down only as a balance-sheet and working-capital adjustment. The benchmark required that omitted FY2025 expense to reduce EBITDA as well.

    The bonus adjustment went wrong in the other direction. DeepSeek treated the $1.2 million retention pool as a separate, wholly unrecorded obligation and deducted it in full. The benchmark's company specification and answer key treat the $600,000 already accrued in payroll as part of that pool, leaving a further $600,000 expense to recognise.

    The two errors partly cancelled out: omitting the inventory expense overstated EBITDA by $720,000, while the excess bonus deduction understated it by $600,000. DeepSeek reported $21.576 million against the expected $21.456 million, both including the conditional $480,000 rent saving. It failed the inventory, bonus and final-total checks. The $120,000 gap in the final figure hid two larger errors that a reviewer would need to correct before relying on the analysis.

    03 / TASK EXPLORERClick a cell to inspect its trials

    50 of 50 tasks · Cells show complete answers out of three.

    TaskDeepSeekSolLunaGLMKimiMiniMax
    001 / easyFY2025 net revenue reconciliation
    002 / easyLargest customer accounts
    003 / easyOverdue receivables
    004 / easyManagement proposed EBITDA adjustments
    005 / easyLargest supplier and contract
    006 / easyMonthly revenue outlier
    007 / easyClosing cash and bank reconciliation
    008 / easyHeadcount and payroll
    009 / easyReported EBITDA
    010 / easyReported gross margins
    011 / easyInventory balances and days
    012 / easyYear-end DSO
    013 / easyCustomer credits
    014 / easyOperating expense categories
    015 / easyActive customer counts
    016 / mediumRevenue growth by customer cohort
    017 / mediumCustomer concentration and control
    018 / mediumManagement accounts and trial balance
    019 / mediumFreight cutoff correction
    020 / mediumMonthly operating working capital
    021 / mediumOverdue customer balance and deal treatment
    022 / mediumSlow stock adjustment
    023 / mediumClosing debt-like items
    024 / mediumRelated-party transactions
    025 / mediumGross-margin bridge
    026 / mediumMonthly revenue pattern and unusual sale
    027 / mediumPayment days and year-end stretch
    028 / mediumSeverance and recurrence
    029 / mediumActual versus budget
    030 / mediumSubsequent credits and cutoff
    031 / mediumKestrel payment terms
    032 / mediumCapex versus depreciation
    033 / mediumYear-end covenant compliance
    034 / mediumBonus correction and net debt
    035 / mediumWarehouse rent normalisation
    036 / hardUnusual December sale
    037 / hardManagement growth and run-rate claims
    038 / hardSupported adjusted EBITDA
    039 / hardManagement margin explanation
    040 / hardEarnings to operating and free cash flow
    041 / hardNormalised working-capital peg
    042 / hardDefensible run-rate EBITDA
    043 / hardKestrel dependence and downside
    044 / hardSupplier economics and margin sustainability
    045 / hardOwner-related normalisations
    046 / hardYear-end presentation and cash timing
    047 / hardSupported covenant headroom
    048 / hardMissing evidence requests
    049 / hardEnterprise to equity bridge
    050 / hardDeal-team red-flag memo

    TASK 019 · medium · GPT-6 Luna

    Freight cutoff correction

    What adjustment to FY2025 EBITDA do year-end cost cutoff issues require?

    Required checkTrial 1Trial 2Trial 3
    Midwest freightPassPassPass
    Lakefront freightPassPassPass
    EBITDA correctionPassPassPass
    Scope and precisionPassPassPass
    Complete answerPassPassPass

    Saved judge verdicts, not an independent human review. Failure of one check fails the task.

    8 tasks pass every answer; 7 never pass. The selected detail remains visible when filtering, so you can keep inspecting a task.

    Resource use

    04 / RESOURCE USEMean per completed trial

    Estimated cost without caching · USD per completed task run

    GPT-6 Luna$0.064
    GLM 5.3 Flash$0.138
    MiniMax M2.5$0.183
    DeepSeek V4.1 Flash$0.881
    GPT-6 Sol$0.945
    Kimi K3$8.982
    Linear bars start at zero. Cost estimates exclude judging and failed attempts, and ignore caching. They are not invoices.

    The estimated cost per completed task run is $0.064 for GPT-6 Luna, $0.138 for GLM 5.3 Flash, $0.881 for DeepSeek V4.1 Flash and $0.945 for GPT-6 Sol, excluding judging and provider caching. Among the original four models, DeepSeek is the least token-efficient: it uses roughly 3 to 7 times as many tokens as the other models and about twice as many turns as Sol.

    Kimi averages 2.63 million tokens and 25.9 turns per completed run, close to DeepSeek’s 2.87 million tokens and 27.5 turns. MiniMax averages 0.59 million tokens and 18.8 turns. Using AWS Bedrock Standard rates for their recorded inference routes, estimated cost per completed run is $8.982 for Kimi and $0.183 for MiniMax.

    AWS lists Kimi K3’s US cross-region rates at $3.30 per million input tokens and $16.50 per million output tokens. MiniMax M2.5’s Oregon rates are $0.30 and $1.20, respectively. These estimates treat input as uncached, matching the original cost comparison; they do not represent actual bills.

    05 / COST EFFICIENCYAll six models · 150 completed trials each

    Higher and further left is better. The dashed Pareto frontier connects models for which no other model is both at least as accurate and at least as cheap, with one strict improvement.

    Task pass rate versus estimated cost per completed runCost uses a logarithmic USD scale. Pass rate uses a linear zero-to-100-percent scale. Frontier: GPT-6 Luna, DeepSeek V4.1 Flash, Kimi K3. Exact values for all six models appear in the table below.0%20%40%60%80%100%$0.05$0.10$0.20$0.50$1$2$5$10Task pass rateEstimated USD per completed run · logarithmic scaleDeepSeek V4.1 Flash: $0.881 per run; 60.0% pass rateDeepSeekGPT-6 Sol: $0.945 per run; 58.7% pass rateSolGPT-6 Luna: $0.064 per run; 50.7% pass rateLunaGLM 5.3 Flash: $0.138 per run; 47.3% pass rateGLMKimi K3: $8.982 per run; 62.0% pass rateKimiMiniMax M2.5: $0.183 per run; 24.7% pass rateMiniMax

    DeepSeek V4.1 Flash: $0.881 per completed run · 60.0% pass rate · 2.849M input tokens + 22.1K output tokens · 27.5 agent turns on average.

    ModelUSD / runTask pass rateUSD / passing answer
    GPT-6 Luna$0.06450.7%$0.126
    GLM 5.3 Flash$0.13847.3%$0.292
    MiniMax M2.5$0.18324.7%$0.744
    DeepSeek V4.1 Flash$0.88160.0%$1.469
    GPT-6 Sol$0.94558.7%$1.610
    Kimi K3$8.98262.0%$14.487
    AWS rates and calculation

    USD per million tokens, verified 2026-09-29. These rates match the inference routes in the saved runs.

    Model / routeInputOutputCache readCache write
    Kimi K3US cross-region · Standard$3.30$16.50$0.33$4.125
    MiniMax M2.5Oregon · Standard on-demand$0.30$1.20Not listedNot listed

    For Kimi and MiniMax: mean cost = (completed-run input tokens × input rate + completed-run output tokens × output rate) / 1,000,000 / 150. Each completed trial counts, whether its answer passes or fails. Superseded attempts, execution failures, judging and infrastructure are excluded.

    All input is priced as uncached for consistency with the original comparison. Kimi supports automatic caching, but the saved usage does not split cache reads and writes, so these are estimates, not invoices. The listed cache rates are not applied; Kimi’s cache-write rate is for a 30-minute retention period.

    The original four models retain the report’s OpenAI Standard and Fireworks serverless pricing snapshots from 24–25 September 2026, including GPT-6 Sol’s long-context pricing adjustment.

    USD per passing answer is total estimated spend on the 150 completed trials divided by passing answers, not a prediction of retry costs. The frontier describes these observed results under uncached pricing. System prompts differed between cohorts, so it does not isolate model differences.

    Under these assumptions, Luna, DeepSeek and Kimi form the observed cost-efficiency frontier. Luna has the lowest estimated cost per run, DeepSeek adds a higher pass rate, and Kimi’s two-point increase over DeepSeek comes at roughly ten times the estimated cost. Sol, GLM and MiniMax each have another model that costs less and passes more often in these results. The different system prompts remain a limitation of this comparison.

    Limitations

    These are first results, not a final ranking of models for financial work.

    All tasks share one company, so they are not independent samples. A failed answer means the response did not satisfy every criterion; it does not tell us how much work a human reviewer would need to fix it. And every criterion is graded by one judge from the same family as one of the evaluated models, calibrated on authored reference and flawed answers rather than practitioner review.

    Tasks 041 and 049 are included using their saved regrades. We had previously excluded them because their questions rely on assumptions that could be stated more clearly; including them does not resolve that concern.

    We plan to address these limitations before making broader claims.

    Conclusion

    The main thing we learned is that deal-side finance is a useful stress test for agents because the work is not just arithmetic. The evidence is spread across formats, the numbers have to reconcile, and the important finding is often something nobody states directly.

    The benchmark only works because the company comes before the documents. Every planted fact has one source of truth, every downstream file is generated from it, and every expected answer can be traced back to the underlying records.

    That also gives us a path forward: build more companies with different stories, not just more questions about Meridian. Next we are expanding the benchmark and adding practitioner review to the grading process.

    Let’s try this on your work.

    Get in touch. Together, we’ll shape how agents handle your team’s work.

    Contact the team

    SecondState

    SecondState provides software for finance teams. Use of SecondState's products and services is subject to our terms of use and privacy policy.

    From question to evidence.
    From evidence to decision.

    Follow the signal wherever it leads, with every conclusion tied back to its source.

    © 2026 SecondState. All rights reserved.