Artificial Analysis、インクリング・スモールがインクリングに迫る性能を低パラメータで達成と報告
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
Artificial Analysisの知能指数によると、インクリング・スモールはパラメータ数が同社の主力モデル「インクリング」の3 分の 1未満でありながら、その性能差がわずか 1 ポイント以内であることが示された。
AI深層分析を開く2026年8月1日 13:57
AI深層分析
キーポイント
高効率なパラメータ設計と性能維持
Thinking Machines は「Inkling Small」において、総パラメータ数276B(アクティブ12B)というサイズで、フラッグシップの「Inkling」とほぼ同等の知能指数40を達成し、同規模以下のオープンウェイトモデルでは最高スコアを更新した。
コーディングと推論タスクでの優位性
Humanity's Last ExamやGPQA Diamondなどの主要な評価基準において、「Inkling Small」はフラッグシップモデルを上回る結果を示し、複雑な推論やコード生成能力に優れた性能を発揮した。
エージェントタスクと事実知識の課題
銀行業務シミュレーションなどのエージェントタスクでは「Inkling」に劣る一方、AA-Omniscience Index において正解数が誤答数を上回らない状況にあるが、これは推論能力の不足よりも事実知識の精度低下によるものである。
出力トークン数の効率性
知能指数評価における平均出力トークン数は約24Kで、「Inkling」や他社競合モデルと比較して少ない計算リソースで同等のタスクを完了できる高い効率が確認された。
Inkling Smallのパラメータ構成と性能
276BパラメータのMoEモデルで、アクティブパラメータは12Bであり、Inklingの3分の1未満である。同サイズ以下のオープンウェイトモデルでは最高スコアを記録し、DeepSeek V4 Flash (max) と同等の結果を出している。
重要な引用
Inkling Small holds a similar tier of intelligence to Inkling's at less than one third of its size: 276B total parameters (12B active) vs. 975B (41B active) for Inkling.
Inkling Small meets or exceeds Inkling on several coding and frontier reasoning evaluations.
No open weights model at its size or smaller scores higher on the Intelligence Index.
Inkling Small is a 276B parameter MoE with 12B active parameters, less than a third of Inkling.
編集コメントを表示
編集コメント
Thinking Machines Lab が公開した「Inkling Small」は、パラメータ数の削減と性能維持の両立という難題に対し、MoE(Mixture of Experts)アーキテクチャを用いた実証的な解決策を示している。この成果は、リソース制約のある環境でのAI活用や、オープンウェイトモデルの競争力向上に寄与する重要な一歩となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
「Thinking Machines」が発表した新モデル「Inkling Small」は、人工知能評価指標「Artificial Analysis Intelligence Index」で40点を記録し、同社の主力モデル「Inkling」と僅か1ポイント差の成績を収めました。驚くべきことに、この性能を実現するために必要なパラメータ数は、主力モデルの3分の1未満に抑えられています。
Thinking Machines Lab が発表した「Inkling Small」は、同社が2週前にリリースした主力モデル「Inkling」に続く2作目のモデルです。サンフランシスコを拠点とするこの AI ラボは、元 OpenAI の CTO であるミラ・ムラティ氏によって設立されました。「Inkling Small」はオープンウェイトの推論モデルで、総パラメータ数は 276B(アクティブな MoE パラメータは 12B)、テキスト・画像・音声の入力を処理でき、コンテキストウィンドウ長は 256K トークンをサポートします。
主な成果:
➤ Inkling Small は、主力モデル「Inkling」と同等の知能レベルを維持しながら、サイズは3分の1未満に抑えています。 総パラメータ数は 276B(アクティブ 12B)に対し、「Inkling」は 975B(アクティブ 41B)です。この規模以下のオープンウェイトモデルで、知能指数のスコアを上回るものは存在しません。同程度の規模を持つ DeepSeek V4 Flash(最大値)も総パラメータ 284B(アクティブ 13B)でスコア 40 を記録していますが、MiniMax-M3 はアクティブパラメータ 23B で 44 点、GLM-5.2(最大値)はアクティブパラメータ 40B で 51 点をそれぞれ達成しています。
➤ Inkling Small は、コーディングや最先端の推論に関する評価において、主力モデル「Inkling」に匹敵するか、あるいはそれを上回る結果を出しました。 「Humanity's Last Exam」では 32%(対 Inkling: 30%)、「GPQA Diamond」では 89%(同:87%)、「CritPt」では 8%(同:5%)、「SciCode」では 49%(同:46%)と、いずれも上回っています。また、「Terminal Bench v2.1」でも 55% と同等のスコアを記録しています。
「Inkling Small」は、エージェントタスクや事実知識の分野では「Inkling」に劣ります。τ³-Banking では 15% と 24% で差がつき、GDPval-AA v2 では 1269 Elo と 1237 Elo でわずかに上回っています。AA-Omniscience Index では -9 点と、フラッグシップモデルの +2 点を下回り、正解よりも誤答の方が多くなっています。これは幻覚(ハルシネーション)によるものではなく、AA-Omniscience の精度が 31% と 40% で低いことが原因です。一方、Inkling Small の幻覚発生率は 57% と、兄弟モデルの 63% よりもわずかに低くなっています。
Inkling Small はインテリジェンス・インデックスの各タスクで平均約 24K トークンの出力を生成しました。これは Inkling の約 25K トークンよりやや少ない数値ですが、同程度の知能レベルを持つ他モデルはさらに多くのトークンを消費しています。DeepSeek V4 Flash は約 45K、GPT-5.4 mini (xhigh) は約 78K です。フルインテリジェンス・インデックスの実行には、Inkling Small と Inkling の両方でほぼ同じ数の出力トークンが必要でした。
追加モデル情報:
- 種類: オープンウェイトの推論モデル(Apache 2.0 ライセンス)
- サイズ: 総パラメータ数 276B、アクティブ 12B (MoE)
- 入力モダリティ: テキスト、画像、音声
- 出力モダリティ: テキスト
- コンテキストウィンドウ: 256K トークン(Inkling は 1M をサポート)
Thinking Machines チームのリリースを祝います!

Inkling Smallは、パラメータ数276BのMoEモデルで、アクティブなパラメータ数は12Bです。これはフルサイズのInklingの3分の1未満に相当します。このサイズまたはそれ以下のオープンウェイトモデルで、これより高いスコアを記録したものは存在しません。DeepSeek V4 Flash(最大値)もアクティブパラメータ数13B(総計284B)で同スコアを達成していますが、MiniMax-M2.7(230B、アクティブ10B)は2ポイント下回っています。

Inkling Smallは、AA-Omniscience(人工知能の全知性指数)においてInklingより低いスコアを記録し、-9点でした。一方、Inklingは2点を獲得しています。これは、正解よりも不正解の方が多いことを意味します。この低スコアの主な要因は精度の低下(31%対40%)であり、総パラメータ数が少ないモデルに典型的な傾向です。ただし、Inkling Smallのハルシネーション率(57%)はInkling(63%)よりわずかに低い水準にあります。

Inkling Smallは、Artificial Analysis Intelligence Indexの各タスクで平均約24Kトークンの出力を生成しました。これはInkling(約25K)よりわずかに少ない数値です。同程度の知能レベルにある他モデルと比較すると、DeepSeek V4 Flash(最大値、約45K)やGPT-5.4 mini(xhigh、約78K)と比べてもトークン効率は非常に良好です。インテリジェンス指数の全評価を実行する際に生成された出力トークンの総数は約131Mで、フラッグシップモデルの約128Mよりわずかに多くなっています。

「Inkling Small」は、人工知能分析インデックス(Artificial Analysis Intelligence Index)において、「Inkling」の上位にランクインしました。パラメータ数は前者が後者の3分の1未満です。
AA-Briefcaseという、長期的なエージェント型知識作業を評価するベンチマークでは、Inkling Small は兄弟モデルである Inkling よりも高いスコアを獲得しています(917 Elo vs. 839)。この差は主にプレゼンテーションの質の高さによるものです。採点基準の通過率は両者ほぼ同じ(20% vs. 19%)であり、Inkling Small の出力が本質的により正確であるというよりは、提示方法が優れていることが示唆されます。ただし、タスク完了までのターン数は大幅に短縮されており、平均して1タスクあたり34ターンで完了したのに対し、Inkling は81ターンを要しました。

人工知能分析インデックスにおける9つの評価項目全体での Inkling Small のパフォーマンス詳細は以下の通りです。

詳細やベンチマークについては、Artificial Analysis の公式サイトをご覧ください:https://artificialanalysis.ai/models/inkling-small
原文を表示
Thinking Machines' new Inkling Small scores 40 on the Artificial Analysis Intelligence Index, within a point of its flagship sibling Inkling with less than a third of the total and active parameters
Inkling Small is Thinking Machines Lab's second model release, arriving two weeks after Inkling launched at 41 on the Artificial Analysis Intelligence Index. Thinking Machines is the San Francisco-based AI lab founded by former OpenAI CTO Mira Murati. Inkling Small is an open weights reasoning model with 276B total parameters (12B active MoE), text, image, and speech input, and a 256K token context window.
Key results:
➤ Inkling Small holds a similar tier of intelligence to Inkling’s at less than one third of its size: 276B total parameters (12B active) vs. 975B (41B active) for Inkling. No open weights model at its size or smaller scores higher on the Intelligence Index. DeepSeek V4 Flash (max), at a similar size of 284B total (13B active) also scores 40, MiniMax-M3 reaches 44 with 23B active, while GLM-5.2 (max) reaches 51 with 40B active.
➤ Inkling Small meets or exceeds Inkling on several coding and frontier reasoning evaluations. It scores higher on Humanity's Last Exam (32% vs. 30%), GPQA Diamond (89% vs. 87%), CritPt (8% vs. 5%), and SciCode (49% vs. 46%), and achieves the same score on Terminal Bench v2.1 (55%).
➤ Inkling Small is not as strong as Inkling on Agentic tasks and factual knowledge. Inkling Small trails Inkling on τ³-Banking (15% vs. 24%), though it edges ahead on GDPval-AA v2 (1269 vs. 1237 Elo). On the AA-Omniscience Index it scores -9 vs. the flagship's positive 2, meaning incorrect answers outweigh correct ones; this is driven by lower AA-Omniscience Accuracy (31% vs. 40%) rather than hallucination: Inkling Small’s Hallucination Rate is slightly lower than its larger sibling (57% vs. 63%).
➤ Inkling Small averaged ~24K output tokens per Intelligence Index task, slightly fewer than Inkling (~25K), while peers at its intelligence level averaged far more. DeepSeek V4 Flash averaged ~45K and GPT-5.4 mini (xhigh) ~78K. Running the full Intelligence Index took roughly the same number of output tokens for Inkling Small and Inkling.
Additional model details:
➤ Type: Open weights reasoning model (Apache 2.0 license)
➤ Size: 276B total parameters, 12B active (MoE)
➤ Input modalities: Text, image, and speech
➤ Output modalities: Text
➤ Context window: 256K tokens (Inkling supports 1M)
Congratulations to the team at Thinking Machines on the release!

Inkling Small is a 276B parameter MoE with 12B active parameters, less than a third of Inkling. No open weights model at its size or smaller scores higher. DeepSeek V4 Flash (max) has 13B active parameters (284B total) and achieves the same score, while MiniMax-M2.7 (230B, 10B active) sits 2 points lower.

Inkling Small scores lower than Inkling on AA-Omniscience, coming in at -9 where Inkling scores 2, meaning it has more incorrect answers than correct ones. The lower score is driven by lower accuracy (31% vs. 40%), which is typical of smaller total parameter count models. Inkling Small's Hallucination Rate is slightly lower (57% vs. 63%).

Inkling Small averaged ~24K output tokens per Artificial Analysis Intelligence Index task, slightly fewer than Inkling (~25K) - token efficient next to peers near its intelligence level like DeepSeek V4 Flash (max, ~45K) and GPT-5.4 mini (xhigh, ~78K). Running the full Intelligence Index evaluations took ~131M output tokens, slightly more than the flagship's ~128M.

Inkling Small reaches a higher score than its larger sibling on AA-Briefcase, our long-horizon agentic knowledge-work benchmark: 917 Elo vs. 839 for Inkling, driven significantly by better presentation. Rubric pass rates are nearly identical (20% vs. 19%), suggesting Inkling Small's outputs are better presented rather than substantially more correct; however, it also finished tasks in less than half the turns (34 vs. 81 per task on average).

Full breakdown of Inkling Small's performance across the nine evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks: https://artificialanalysis.ai/models/inkling-small
AI算出
主要ニュースainew評価標準
AI モデルの性能比較、パラメータ効率、および詳細なベンチマークスコアを報じており、新規性の高い主要ニュースとして分類される。日本企業との直接的な関連性は薄いものの、オープンウェイトモデルとしての技術的価値は高い。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み