会計分野の AI 生産性ベンチマーク「APEX-Accounting」が公開
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
AI による会計業務の生産性を評価するためのベンチマーク「APEX-Accounting」が発表され、6 分程度の読み物として提供された。
AI深層分析を開く2026年8月3日 23:51
AI深層分析
キーポイント
実務レベルでのAI性能限界
8回の試行で全てのタスクを正解できるモデルは存在せず、最良のモデルでも成功率は2.6%に留まった。
ベンチマークの厳格な設計
デロイトやPwCなどの経験者が作成した10社分のシミュレーションデータを用い、単なる正解ではなく一貫性を重視して評価を行った。
モデルごとの性能差と特性
Claude Fable 5 が平均 56.4% で首位だが、予算に対する感応度には大きな違いがあり、Muse Spark 1.1 は低予算でも安定している。
試験と実務の乖離
試験では正解を出せるモデルも、実際のワークフローにおける判断や文脈の維持が求められれば失敗するケースが多発した。
予算超過後の性能飽和
$10 から$50 に予算を増やしても使用トークンは平均64.7%に留まり、モデルは最大予算を十分に活用しない。
重要な引用
Even the most consistent model solved just 2.6% of tasks correctly in all eight runs.
This is a very different result from performance on accounting exams.
58% of tasks were never fully solved by any model on any run.
Roughly seven in ten failures came from flawed reasoning, not an inability to find the right information.
編集コメントを表示
編集コメント
このベンチマークは、AI が「正解を導き出す能力」と「業務を完遂する能力」の間に大きな隔たりがあることを浮き彫りにした。業界全体が実用化を目指す中で、この結果は過剰な期待を抑制し、現実的な導入ロードマップの策定に重要な示唆を与えるものだ。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Ramp との共同開発により、新しいベンチマーク「APEX-Accounting」が発表されました。
会計分野のベンチマークは往々にして、モデルが一度だけ正解を導き出せるかどうかを試すものになりがちです。しかし、決算処理にはそれ以上の能力が求められます。エージェントは矛盾するファイルの整合性を取ったり、企業固有の文脈を適用したり、結論を複数のステップにわたって維持しながら正確な結果を一貫して出力する必要があります。中間段階で正解を出しても、最終的な仕訳が間違っていれば意味がありません。
APEX-Accounting は、Ramp が会計士向けの AI オペレーティングプラットフォーム「Stack」を開発する過程で構築した評価システムを基盤としています。Ramp と Mercor は、AI エージェントが実際の会計業務をプロの水準で完遂できるかを検証するためにこのベンチマークを共同開発しました。
対象となるのは 10 の架空企業における 160 のタスクです。これらのタスクと評価基準は、Mercor のプラットフォームを通じて募集された、デロイト、PwC、EY、KPMG などの大手監査法人で実務経験を持つ専門家によって作成されました。
注目すべきは、どのモデルがトップを走るかではなく、いかに多くのモデルが一貫して成功できないかという点です。会計業務では一度の正解ではなく、繰り返し正確であることが求められるため、すべてのモデルを各タスクで 8 回実行しました。最も安定したモデルでも、8 回の試行すべてで正解できたのは全体の 2.6% に過ぎませんでした。
これは会計試験におけるパフォーマンスとは全く異なる結果です。制御されたテスト環境では正解を出せても、実際のワークフローを通じて会計上の判断を維持できなければ、現場では失敗することになります。
リーダーボードの結果
Claude Fable 5 が 56.4% でトップに立ち、次いで Meta の Muse Spark 1.1 が 52.6%、GPT-5.6 Sol が 51.5% です。最上位のモデルでも、専門家がこなす作業をわずかに半分超える程度しか完了できていません。
8 回の試行のうち少なくとも 1 回で成功したかどうかを測る「Pass@8」では、Muse Spark 1.1 が 21.5% で最高となり、Fable 5(20.1%)をわずかに上回っています。
多くのモデルが部分的な正解として評価されました。少なくとも 1 つのモデルが、タスクの 95% 以上で何らかの評価を得ています。ただし完全な解答は難しく、どのモデルもどのランでも 58% のタスクは一度も完全に解決できませんでした。

パフォーマンスとコスト
タスクあたりの予算を 1 ドル、5 ドル、10 ドル、そして 50 ドルに設定してモデルをテストしました。予算を増やすことで利用可能なトークン数が増え、一般的には結果が向上します。しかし、その効果はモデルによって大きく異なります。
Fable 5 は予算に対して極めて敏感でした。1 ドルの予算ではスコアが 11.8% でしたが、50 ドルに引き上げると 55.2% に改善しています。一方、Muse Spark 1.1 はその逆です。厳しい予算制約下でもすでに高い性能を発揮しており、予算を増やしてもほとんど向上しません。
これは、1 ドルの上限を超えるとモデルがトークン制限を十分に活用できていないためです。10 ドルから 50 ドルへ予算を引き上げてもトークン使用量は中程度の増加にとどまり、平均するとモデルは 50 ドルという最大予算の約 64.7% しか利用していません。
各モデルが上限付きの予算をどう活用するかには違いがあります。最大予算である 50 ドルの場合でも、Fable 5 は実行あたり約 32 ドルを使用する一方、Muse Spark 1.1 は約 5 ドルです。にもかかわらず、両者のスコア差はわずか 4 ポイント以内です。

リーダーボードが測定するもの
このリーダーボードでは、Mercor の標準的なモデルおよびツールアーキテクチャである「Loop Harness」を用いて全モデルを評価します。また、論文内では、専門的な会計エージェントのアーキテクチャの一部を模倣するために Ramp と共同で構築された専用「Ramp Harness」との比較も別途行っています。その結果、Ramp Harness を使用すると、平均して Mean Criteria@3 が 1.2 ポイント向上することが確認されました。
なお、これは Ramp Stack の完全な製品機能に対する評価ではなく、ハイスループットテスト(harness)におけるアブレーション実験の一環です。そのため、Stack の本番環境での統合機能や、会計スキル、メモリ管理、統制機能、監査可能性、あるいは最新のスプレッドシートツールなどについては測定していません。
モデルが失敗する理由
会計および簿記の専門家と協力し、タスク完了を阻むミスの分類体系(タクソノミー)を作成しました。そこには、推論、情報収集、指示への従順さ、そして計画と振り返りにおける失敗が含まれます。
上位 3 つのモデルはいずれも同様の理由で失敗しています。失敗の約 7 割は、適切な情報を取得できないことではなく、推論に欠陥があることに起因します。例えば、ワークフローの初期段階で差異を正しく特定しながらも、最終的な仕訳においてその発見を省略したり、矛盾した記述をしたりするケースが典型です。単に情報検索能力を向上させるだけではこの問題は解決しません。モデルには、より優れた会計判断力と、結論をワークフロー全体を通じて一貫して維持するための規律が必要です。

方法論
このベンチマークの課題作成と解決には、会計専門家 40 人以上が携わりました。参加者の経験年数の中央値は 11 年で、半数以上が四大監査法人で勤務した経験があります。
各タスクは、月次決算直前の状態に固定された架空の会社という独立した世界の中で行われます。そこには独自の勘定科目、記録、事業履歴、会計ソフトウェア、スプレッドシート、PDF、その他のファイルが存在します。
企業自体は架空ですが、そこに含まれる取引、不整合、そしてエッジケースはすべて、実務でこの仕事に携わる会計士によって作成されました。各タスクには専門家による評価基準が設けられており、平均 13.7 の項目が含まれています。評価はオープンソースの AI ジャッジが行い、専門家の採点者との一致率は 97% に達しています。
APEX-Accounting は、月次決算処理と記帳ワークフローに焦点を当てています。税務申告、監査、連結会計、多拠点・多通貨対応の会計、外部への報告、あるいは不明点が生じた際のエージェントの対応については評価対象外です。本稿で得られた知見は、この範囲内で理解する必要があります。
APEX-Accounting リーダーボードには、10 の世界に属する n=160 件のタスクからなるホールドアウトテストセットが採用されています。データ汚染を防ぐため、これらのデータは非公開のままです。一方、Hugging Face では、世界 1 つ・タスク 10 件のオープンソースサンプルセットを公開しています。また、エージェント評価を実行するための内部フレームワーク「Archipelago」は GitHub で利用可能です。
詳細な手法と結果については、技術報告書 をご参照ください。
APEX ベンチマーク
APEX-Accounting は、経済的に価値のある業務を AI モデルが遂行できる能力を評価する Mercor 社の APEX ベンチマークシリーズに新たに加わりました。同シリーズには、統合と観測性を伴うソフトウェアエンジニアリングタスクを対象とした「APEX-SWE」、企業法務・経営コンサルティング・投資銀行といった専門サービス分野を対象とした「APEX-Agents」が含まれます。
APEX-Accounting の作成に時間を割いてくださった Ramp 社および Mercor マーケットプレイス上のすべての会計士の方々に、心より感謝申し上げます。
APEX-Accounting はクローズドなベンチマークであるため、リクエストに応じて最先端モデルに対するリーダーボード評価の実行が可能です。詳細については こちら からお問い合わせください。
原文を表示
Introducing APEX-Accounting, a new benchmark built in collaboration with Ramp.
Accounting benchmarks often test whether a model can produce the right answer once. Closing the books demands more. An agent must reconcile conflicting files, apply company-specific context, carry conclusions across multiple steps, and produce the correct result consistently. A model that drops a correct intermediate answer can still create a bad journal entry. APEX-Accounting builds on the evaluation system Ramp developed while building Stack, their AI operating platform for accountants. Ramp and Mercor built APEX-Accounting to test whether AI agents can complete real accounting work to professional standards. The benchmark includes 160 tasks across 10 simulated companies. Experts with prior experience at firms like Deloitte, PwC, EY, and KPMG hired through Mercor's platform created the tasks and grading rubrics.
The headline result is not which model leads, but how rarely any model succeeds consistently. We ran every model on every task eight times because accounting work must be repeatedly correct, not correct once. Even the most consistent model solved just 2.6% of tasks correctly in all eight runs.
This is a very different result from performance on accounting exams. Models can produce the right answer in a controlled test and still fail to carry accounting judgment through a real workflow.
Leaderboard results
Claude Fable 5 tops the leaderboard at 56.4%, followed by Meta's Muse Spark 1.1 at 52.6%, and GPT-5.6 Sol at 51.5%. The best models manage to complete just over half of the work that a professional would.
For Pass@8, which measures whether a model is successful in at least one of eight attempts, Muse Spark 1.1 is highest at 21.5%, just ahead of Fable 5 (20.1%).
Models frequently earned partial credit. At least one model earned some credit on more than 95% of tasks. Full solutions were more difficult. 58% of tasks were never fully solved by any model on any run.

Performance vs. cost
We tested models at spending budgets of $1, $5, $10, and $50 per task. A larger budget allows models to use more tokens, which generally improves results. But the effect varies enormously. Fable 5 was extremely budget sensitive, scoring 11.8% with a $1 budget and improving to 55.2% with a $50 budget. Muse Spark 1.1 is the opposite. It's already strong on a tight budget and barely improves with more money. This is because above the $1 cap, models typically do not fully utilize the token limits. There is only a moderate increase in token usage when going from $10 to $50 and, on average, models use just 64.7% of the maximum budget available at $50.
Models differ in how they utilize the capped budget. At the $50 maximum spending budget, Fable 5 actually spends ~$32 per run while Muse Spark 1.1 spends ~$5, yet their scores are within 4 percentage points.

What the leaderboard measures
The canonical leaderboard runs every model in the Loop Harness, Mercor's standard model-and-tools architecture. The paper separately compares it with a purpose-built Ramp Harness, created with Ramp to mirror parts of a specialized accounting-agent architecture. The Ramp Harness improves Mean Criteria@3 by 1.2 percentage points on average.
This is a harness ablation, not an evaluation of the complete Ramp Stack product. It does not measure Stack's production integrations, accounting skills, memory, controls, auditability, or recent spreadsheet tooling.
Why do models fail?
Working with accounting and bookkeeping experts, we developed a taxonomy of the mistakes that prevented models from completing each task. These included failures in reasoning, information gathering, instruction following, and planning and reflection.
The top three models all fail in much the same way. Roughly seven in ten failures came from flawed reasoning, not an inability to find the right information. A model might correctly identify a discrepancy early in the workflow, then omit or contradict that finding in its final journal entry. Better retrieval alone will not solve this. Models need stronger accounting judgment and more discipline in carrying conclusions through an entire workflow.

Methodology
More than 40 accounting professionals authored and solved the benchmark tasks. They had a median of 11 years of experience, and more than half had worked at a Big Four accounting firm. Each task takes place in a self-contained world: a fictional company frozen at month-end close, with its own accounts, records, business history, accounting software, spreadsheets, PDFs, and other files.
The companies themselves are fictional but every transaction, discrepancy, and edge case inside was written by accountants who do this work for a living. The experts authored a rubric for each task with, on average, 13.7 criteria. An open-source AI judge performs the grading and achieves 97% agreement with expert graders.
APEX-Accounting focuses on month-end close and bookkeeping workflows. It does not evaluate tax, audit, consolidation, multi-entity or multi-currency accounting, external reporting, or how agents respond when they need clarification. Its findings should be understood within that scope.
The APEX-Accounting leaderboard comprises a heldout test set of n=160 tasks (associated with 10 worlds), kept hidden to resist contamination. We have released an open-source sample set with 1 world and 10 tasks on Hugging Face, and Archipelago, our internal framework for running agent evaluations, is available on GitHub.
Read the APEX-Accounting technical report to learn more about the methodology and results.
APEX benchmarks
APEX-Accounting joins Mercor's family of APEX benchmarks, which evaluate AI models' ability to do economically valuable work. Other APEX benchmarks include APEX-SWE for software engineering tasks across integration and observability, and APEX-Agents for professional services domains including corporate law, management consulting, and investment banking.
We thank Ramp and all the accountants on the Mercor marketplace who contributed their time to creating APEX-Accounting.
As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request. Reach out here to find out more.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み