Thinking Machines、米国製オープンウェイトモデル「Inkling」を公開
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
Thinking Machines が新モデル「Inkling」を公開し、人工知能指数で米国製オープンウェイトモデル首位となるなど、エージェント性能やマルチモーダル処理において顕著な成果を示した。
AI深層分析を開く2026年8月1日 14:14
AI深層分析
キーポイント
米国製オープンウェイトモデルの首位獲得
Inkling は Artificial Analysis の知能指数で 41 を記録し、前作である Nemotron 3 Ultra や Gemma 4 などを上回り、米国ラボからの最新オープンウェイトモデルとして最高位にランクインした。
エージェント性能における他社モデルの凌駕
GDPval-AA v2 および𝜏³-Banking のベンチマークにおいて、Inkling は Kimi K2.6 や DeepSeek v4 Flash を上回るスコアを記録し、特にタスク遂行能力に優れていることが示された。
トークン効率とマルチモーダル対応の強化
Inkling は出力トークンの平均使用量が競合他社より少なく効率的であり、テキストに加え画像や音声入力をネイティブで処理できる点でオープンウェイトモデル間の差別化要因となっている。
大規模パラメータとコンテキストウィンドウの提供
同モデルは総パラメータ数 975B(アクティブ 41B)を有し、Tinker API では 256K、HuggingFace のウェイト公開版では 1M という広範なコンテキストウィンドウをサポートしている。
GDPval-AA v2での高いスコア
InklingはElo 1238を記録し、Kimi K2.6やDeepSeek v4 Flash maxを上回る性能を示した。
重要な引用
Thinking Machines has released Inkling, the new leading U.S. open weights model
Inkling debuts at 41 on the Artificial Analysis Intelligence Index, making it the leading open weights release from a U.S. lab.
Inkling stands out on agentic performance.
Inkling scores an Elo of 1238 on GDPval-AA v2, higher than Kimi K2.6 (1190) and DeepSeek v4 Flash max (1189)
編集コメントを表示
編集コメント
Thinking Machines の新モデルは、オープンウェイト領域における米国勢の技術的復権を示す重要な指標となっている。特にエージェント機能とマルチモーダル処理の両立において実用性が高いことから、開発コミュニティでの注目度はさらに高まると予想される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Thinking Machines が、米国のオープンウェイトモデルとして新たなトップに君臨する「Inkling」をリリースしました。Artificial Analysis のインテリジェンス指数では 41 位でデビューしています。
Thinking Machines はこれまで研究用のプレビューモデルを発表してきましたが、これは同社初の商用言語モデルです。このモデルは総パラメータ数が 975B、アクティブパラメータ数は 41B を誇り、テキスト・画像・音声の多様な入力に対応しています。
利用方法は、Thinking Machines の「Tinker」プラットフォーム API(コンテキストウィンドウ 256K)を通じてアクセス可能。また、重みデータは HuggingFace で公開されており、こちらはコンテキストウィンドウ 1M をサポートします。
主な成果:
➤ Inkling は Artificial Analysis のインテリジェンス指数で 41 点を獲得し、米国のラボから登場したオープンウェイトモデルの中で最高位となりました。 インテリジェンス指数では前回の米国トップである Nemotron 3 Ultra(38 点)を 3 ポイント上回り、Gemma 4 31B(29 点)や gpt-oss-120b(24 点)も凌駕しています。
➤ エージェント性能において特に目覚ましい成果を残しました。 GDPval-AA v2 と 𝜏³-Banking の両テストで、Kimi K2.6 や DeepSeek v4 Flash を上回るスコアを記録。GDPval-AA v2 では Elo 1238 を獲得し、Kimi K2.6(1190)や DeepSeek v4 Flash max(1189)を上回りました。また 𝜏³-Banking では 24% の正答率を達成し、Kimi K2.6(21%)を上回り、DeepSeek v4 Flash max(23%)にもわずかに及んでいます。
「Inkling」は、オープンウェイトモデルの他社製品と比較してトークン効率に優れています。インテリジェンス・インデックスのタスクにおける出力トークンの平均数は 25K で、GLM-5.2(最大値)の 43K、Kimi K2.6 の 38K、DeepSeek v4 Pro(最大値)の 37K を上回っています。
また、「Inkling」は画像と音声のマルチモーダル入力をネイティブでサポートしており、これがオープンウェイトモデル間の重要な差別化要因となっています。テキスト、画像、音声をすべて入力として受け付けます。画像や動画は階層的なパッチエンコーダーによって符号化され、音声は離散トークン符号化が行われます。これらすべてのモダリティは共有された隠れ空間に投影され、デコーダーによって同時に処理されます。
追加的なモデル詳細:
- サイズ: 975B(アクティブパラメータ 41B)
- 入力モダリティ: テキスト、画像、音声(出力はテキストのみ)
- コンテキストウィンドウ: Tinker では 256K トークン。オープンウェイトモデル版では最大 1M をサポート
- 料金(64K コンテキストウィンドウ、100 万トークンあたり): 入力 $1.87 / キャッシュ $0.374 / 出力 $4.68
- 料金(256K コンテキストウィンドウ、100 万トークンあたり): 入力 $3.74 / キャッシュ $0.748 / 出力 $9.36

GDPval-AA v2 における Elo スコアは 1238 で、Kimi K2.6(1190)や DeepSeek v4 Flash max(1189)を上回っています。

Inkling は、オープンウェイトモデルのリーダーと比較してトークン効率が優れており、インテリジェンスインデックスのタスクあたりの平均出力トークン数は 25K です。これに対し、GLM-5.2(最大値)は 43K、Kimi K2.6 は 38K、DeepSeek v4 Pro(最大値)は 37K を記録しています。

Inkling の AA-Omniscience スコアは +2 で、主要なオープンウェイトモデルには及びませんが、他の米国製オープンウェイトモデルよりは上回っています。次点の Nemotron 3 Ultra は -1 です。精度(Accuracy)スコアは 40% ですが、ハルシネーション率(Hallucination Rate)では 63% を記録しています。

Inkling のパフォーマンスの詳細な内訳は以下の通りです。

詳細なベンチマークについては、Artificial Analysis をご覧ください。
原文を表示
Thinking Machines has released Inkling, the new leading U.S. open weights model, debuting at 41 on the Artificial Analysis Intelligence Index
Thinking Machines has previously released research previews of models and this is their first production language model release. The model is 975B total parameters, has 41B active parameters, and accepts text, image, and audio input modalities. The model is accessible via Thinking Machines’ Tinker platform API (256K context window) and weights are available on HuggingFace (1M context window).
Key results:
➤ Inkling debuts at 41 on the Artificial Analysis Intelligence Index, making it the leading open weights release from a U.S. lab. Inkling scores 3 points higher on the Intelligence Index (41) than the previous leading U.S. open weights model, Nemotron 3 Ultra (38), and also beats Gemma 4 31B (29) and gpt-oss-120b (24)
➤ Inkling stands out on agentic performance. It scores higher than both Kimi K2.6 and DeepSeek v4 Flash on both GDPval-AA v2 and 𝜏³-Banking: Inkling scores an Elo of 1238 on GDPval-AA v2, higher than Kimi K2.6 (1190) and DeepSeek v4 Flash max (1189) and scores 24% on 𝜏³-Banking, higher than Kimi K2.6 (21%) and just above DeepSeek v4 Flash max (23%)
➤ Inkling is token efficient compared to open weights leaders. Inkling averages 25K output tokens per Intelligence Index task compared to 43K, 38K and 37K by GLM-5.2 (max), Kimi K2.6 and DeepSeek v4 Pro (max) respectively
➤ Inkling natively supports image and audio multimodal inputs, a key differentiator among open weights models. Inkling accepts text, image, and audio input modalities. Images and videos are encoded via a hierarchical patch encoder and audio via discrete token encoding, with all modalities projected into a shared hidden space and processed jointly by the decoder
Additional model details:
➤ Size: 975B (41B active) parameters
➤ Input modalities: Text, image, and audio (text output)
➤ Context window: 256K tokens on Tinker, open weights model supports 1M
➤ Pricing per 1M tokens (64K context window): $1.87 input / $0.374 cached / $4.68 output
➤ Pricing per 1M tokens (256K context window): $3.74 input / $0.748 cached / $9.36 output

Inkling scores an Elo of 1238 on GDPval-AA v2, higher than Kimi K2.6 (1190) and DeepSeek v4 Flash max (1189)

Inkling is token efficient compared to open weights leaders, averaging 25K output tokens per Intelligence Index task compared to 43K, 38K and 37K by GLM-5.2 (max), Kimi K2.6 and DeepSeek v4 Pro (max) respectively

Inkling scores +2 on AA-Omniscience, below leading open weights models but above other U.S. open weights models, with the next best Nemotron 3 Ultra (-1). Inkling scores 40% on Accuracy but 63% on the Hallucination Rate

Full breakdown of Inkling's performance:

See Artificial Analysis for further details and benchmarks:
AI算出
主要ニュースainew評価標準
Thinking Machines が初の商用モデルとして Inkling を公開し、既存のトップモデルを上回る性能を示したという明確な新規性があり、AI モデル発表の核心記事である。ただし、日本企業への直接的な影響や日本語一次情報の記載はないため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み