Thinking Machines Lab の Inkling、エージェント型知識作業での性能を検証
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Inkling は AA-Briefcase ベンチマークで Elo 836 を獲得し、DeepSeek V4 Flash よりも高いが、上位オープンモデルには及ばない。
AI深層分析を開く2026年8月1日 14:06
AI深層分析
キーポイント
AA-Briefcase ベンチマークでの評価
Thinking Machines Lab の Inkling は「AA-Briefcase」で Elo 836 を記録し、DeepSeek V4 Flash より上位だが、Nemotron 3 Ultra や GLM-5.2 などの主要オープンウェイトモデルには劣る結果となった。
非標準ファイル処理の弱点
Inkling はマルチモーダル対応を有するものの、Excel、PowerPoint、PDF、Word 以外の非標準ファイルを含むタスクにおいて、正答率(Rubric)が最も低くなる傾向を示した。
プレゼンテーションと分析品質の乖離
Inkling はプレゼンテーション品質(Elo 863)で分析品質(Elo 764)を上回っており、専門的な構成よりも視覚的・形式的な完成度において高い評価を得ている。
リソース消費とターン数の特徴
1 タスクあたりの平均出力トークン数は 52K で同スコア帯のモデルよりやや多く、Excel 生成に最も多くのリソースを要する一方、平均ターン数(81)は高いがツール呼び出し頻度は低い。
AA-Briefcase ルーブリックでのスコア
Inkling は AA-Briefcase ルーブリックで 19.3% を記録し、MiMo-V2.5-Pro に劣るものの DeepSeek V4 Flash max や Gemini 3.5 Flash-Lite よりも高い。
重要な引用
Thinking Machines Lab's Inkling scores an Elo of 836 on on our agentic knowledge work benchmark AA-Briefcase, ahead of DeepSeek V4 Flash but below leading open weights models including Nemotron 3 Ultra and GLM-5.2
Inkling's rubric score is worst on tasks which include non standard 'Other' file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support
Performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively
Inkling scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8)
編集コメントを表示
編集コメント
今回の評価は、単なる言語理解能力だけでなく、実際の業務タスク(文書作成やデータ処理)を完遂する「エージェント」としての能力を厳しく問うものであり、実用化に向けた重要な指針となる。特に非標準ファイルへの対応や分析深度における課題は、開発者が今後のモデル改善で注力すべき領域を示唆している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
モデルページはこちら:https://artificialanalysis.ai/models/inkling
Thinking Machines Lab の「Inkling」は、エージェント型知識作業における評価基準 AA-Briefcase で Elo 836 を記録しました。これは DeepSeek V4 Flash よりも上ですが、Nemotron 3 Ultra や GLM-5.2 といった主要なオープンウェイトモデルには及びません。
新たに導入したエージェント型知識作業用ベンチマーク「AA-Briefcase」は、スプレッドシートやプレゼン資料、UI のモックアップなど、実際の業務で求められる成果物を生成するタスクを数千もの入力ファイルを通じてテストします。モデルの性能は、3 つの観点から評価されます。まず、正解との一致度を判定する二値基準チェックです。次に、分析の質を比較対照して評価する pairwise grading(ペア比較評価)と、プレゼンテーションの質を同様に評価する pairwise grading です。AA-Briefcase の Elo スコアは、これら 3 つの次元の結果を統合した単一の指標となります。
主なポイント:
➤ AA-Briefcase の基準スコアは 19.3%。 これは MiMo-V2.5-Pro(21.4%)には劣りますが、DeepSeek V4 Flash max(18.7%)や Gemini 3.5 Flash-Lite(14.8%)を上回っています。ただし、Inkling の基準スコアが最も低かったのは、Excel、PowerPoint、PDF、Word といった標準的なファイル形式ではなく、「Other」として分類される非標準ファイルタイプを含むタスクでした。これは、同モデルがネイティブにマルチモーダル対応しているにもかかわらずです。
「プレゼンテーションの質」における評価は「分析の質」よりも高く、それぞれエロスコアで 863 と 764 を記録しました。プレゼンテーションと分析の質は、モデル提出物から独立して行われるペア比較チェックによって測定されます。評価者は同じタスクに対する 2 つの提出物を比較し、よりプロフェッショナルに仕上げられている方(プレゼンテーション)と、より深く構造化された分析がなされている方(分析の質)を選出します。
AA-Briefcase タスク 1 つあたり平均で 52K トークンを出力し、フルセット全体では 5M トークンを消費します。これは同程度のスコアを持つ他モデルと比べてやや多めです。特に Excel の成果物に対して最も多くのトークンを使用しており、次いで Word、PDF、PowerPoint、その他順となっています。
AA-Briefcase タスクあたりの平均ターン数は 81 と非常に高く、その変動幅も広いため、中央値は 49 と低くなっています。タスクあたりの平均ターン数が高いにもかかわらず、インクリングは 1 ターンあたりのツール呼び出し回数が平均 0.5 と比較的少ないのが特徴です。

インクリングの AA-Briefcase ルーブリックでのスコアは 19.3% です。これは MiMo-V2.5-Pro(21.4%)には及びませんが、DeepSeek V4 Flash max(18.7%)や Gemini 3.5 Flash-Lite(14.8%)を上回っています。

Inkling の評価スコアは、プレゼンテーション品質の方が分析品質よりも高く、それぞれ Elo スコアで 863 と 764 を記録しています。

Inkling の評価基準では、Excel、PowerPoint、PDF、Word といった標準的なファイル形式ではなく、それら以外の非標準ファイル("Other" ファイル)を含むタスクにおいて、スコアが最も低くなる傾向があります。これは、同モデルがマルチモーダル対応を備えているにもかかわらずです。

Inkling は、AA-Briefcase タスク 1 つあたり平均で 52K トークンの出力を使用し、フルスイート全体では 5M トークンに達します。これは、同様のスコアを持つ他モデルと比べてやや多めです。使用されるトークン数は、Excel の成果物が最も多く、次いで Word、PDF、PowerPoint、そしてその他のファイル形式の順となっています。

Inkling は、AA-Briefcase タスクあたりの平均ターン数が 81 と非常に高く、かつその範囲も広いため、中央値は 49 と大幅に低くなっています。

タスクあたりの平均ターン数が比較的高いにもかかわらず、Inkling は 1 ターンあたりのツール呼び出し回数は平均で 0.5 と、他モデルと比べて少ない傾向にあります。

詳細は以下のリンクをご覧ください。
原文を表示
Thinking Machines Lab’s Inkling scores an Elo of 836 on on our agentic knowledge work benchmark AA-Briefcase, ahead of DeepSeek V4 Flash but below leading open weights models including Nemotron 3 Ultra and GLM-5.2
Our new agentic knowledge work benchmark, AA-Briefcase, tests models on realistic tasks across thousands of input files, requiring deliverables such as spreadsheets, presentations, and UI mock-ups. Model performance is measured across three dimensions: binary rubric checks for ground-truth correctness, pairwise grading on analytical quality, and pairwise grading on presentation quality. The AA-Briefcase Elo is a single metric that combines results across all three dimensions
要点
➤ Scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%). Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support
➤ Performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively. Presentation and Analytical Quality are both measured using separate, independent pairwise checks from model submissions. Graders compare two submissions for the same task and pick the one that is more professionally presented (Presentation) and the one with deeper, better-structured analysis (Analytical Quality)
➤ Uses 52K output tokens on average per AA-Briefcase task and 5M output tokens for the full suite, slightly more than models with similar scores on AA-Briefcase. Inkling uses the most tokens for Excel deliverables, followed by Word, PDF, PowerPoint, and Other types
➤ Has one of the highest mean turns per AA-Briefcase task (81) with also one of the widest ranges, resulting in a much lower median (49). Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5)

Inkling scores 19.3% on the AA-Briefcase rubric, below MiMo-V2.5-Pro (21.4%) but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%)

Inkling performs higher on Presentation than Analytical Quality with Elo scores of 863 and 764 respectively

Inkling’s rubric score is worst on tasks which include non standard “Other” file types (i.e., not Excel, PowerPoint, PDF, or Word), despite its native multimodal support

Inkling uses 52K output tokens on average per AA-Briefcase task and 5M output tokens for the full suite, slightly more than models with similar scores on AA-Briefcase. Inkling uses the most tokens for Excel deliverables, followed by Word, PDF, PowerPoint, and Other types

Inkling has one of the highest mean turns per AA-Briefcase task (81) with also one of the widest ranges, resulting in a much lower median (49)

Despite one of the higher average turns per task, Inkling comparatively uses fewer tool calls per turn on average (0.5)

For more details:
AI算出
技術分析ainew評価標準
Thinking Machines Lab の Inkling という特定の AI エージェントが、新しい独自ベンチマークでどのように振る舞うかを定量的に評価した記事であり、AI モデルの実装・運用に関する具体的な技術的知見を提供している。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み