韓国 AI ラボ Upstage、推論モデル Solar Pro 4 を公開し知能指数が大幅向上
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Artificial Analysis
韓国AI研究所Upstageが推論モデル「Solar Pro 4」をリリースし、知能指数が14から42へ大幅に向上したが、遅延時間の増加とハルシネーション率の上昇というトレードオフが生じた。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 02:05
AI深層分析
キーポイント
知能指数の劇的な上昇
Solar Pro 4はArtificial Analysis Intelligence Indexで42点を記録し、前モデルの14点から27ポイント増加した。
エージェント機能と長期コンテキストの強化
Terminal-BenchやAA-LCRなどのベンチマークで大幅な改善が見られ、特に実世界のエージェントタスクにおけるスコアは人間基準を大きく上回った。
ハルシネーション率と遅延時間の増加
質問への回答試行率が低下したことでオムニセンス評価が改善したが、ハルシネーション率は前モデルから大幅に悪化し、処理速度も低下している。
価格設定とリソース効率
1Mトークンあたりの入出力価格が倍増したが、出力トークン数は減少しており、知能指数に対するコストパフォーマンスは前モデルより向上した。
実世界の自律的作業における大幅な改善
Solar Pro 4 は GDPval-AA v2 で Elo 1277 を記録し、前作の 498 から劇的に向上した。このスコアは人間の基準値(1000)を上回り、Qwen3.7 Max や MiMo-V2.5-Pro と同等以上の性能を示している。
重要な引用
Solar Pro 4 is Upstage AI's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026.
AA-Omniscience improvement from -53 to -1 was from abstaining on more questions.
The intelligence gain comes with a hit to latency.
Solar Pro 4's strongest improvement is on real-world agentic work.
編集コメントを表示
編集コメント
Solar Pro 4は推論能力の劇的な向上を示したが、速度とハルシネーション率という課題を同時に抱えている。このモデルは高度なタスク処理に適しているが、実装時にはコストとパフォーマンスのバランスを慎重に検討する必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
韓国 AI ラボ「Upstage」が Solar Pro 4 を発表。人工知能分析指数(Artificial Analysis Intelligence Index)で 42 を記録し、Solar Pro 3 の 14 から大幅に向上
Upstage AI が新たに開発した独自フラグシップ推論モデル「Solar Pro 4」が、2026 年 4 月に登場した Solar Pro 3 に代わり発表されました。人工知能分析指数では 42 を記録し、「Inkling (xhigh)」(42) と並び、「MiMo-V2.5-Pro」(43) の直後に位置します。これは Solar Pro 3 よりも 27 ポイントの向上です。
Upstage の公式 API を通じた利用料金は、Solar Pro 3 の入力/出力/キャッシュヒット 100 万トークンあたりそれぞれ $0.15/$0.60/$0.02 から、$0.30/$1.20/$0.06 に引き上げられました。
主な結果
➤ Solar Pro 4 が Solar Pro 3 よりも最も大きく改善したのは、エージェント機能と長文コンテキストの処理です。 Terminal-Bench v2.1 では 12% から 57% に、AA-LCR では 31% から 71% に、τ³-Banking では 9% から 23% にそれぞれ向上しました。GDPval-AA v2 における実世界のエージェントタスクでも大きな進歩が見られ、Solar Pro 3 が人間基準(Elo 1000)を大きく下回る 498 のスコアだったのに対し、Solar Pro 4 は 1277 を記録しています。
➤ AA-Omniscience のスコアが -53 から -1 に改善したのは、回答を控えるケースが増えたためです。 Solar Pro 4 は質問への回答を試みる割合が 41% であり、Solar Pro 3 の 92% と比べて大幅に低下しています。その一方で、ハルシネーション(幻覚)発生率は 24% で、Command A+ (14%) や MiniMax-M3 (18%) よりも高いものの、Solar Pro 3 の 88% からすれば劇的な改善です。AA-Omniscience の精度自体は 19% のまま変わっていません。
Solar Pro 4 は、その知能レベルに対してまだ冗長な部分はあるものの、Solar Pro 3 よりもトークン効率が向上しています。Intelligence Index のタスクあたりの出力トークンは 4.3 万で、Solar Pro 3 の 5.2 万(約 17% 減)を下回っています。
知能の向上にはレイテンシの増加が伴います。出力トークン数が削減されているにもかかわらず、平均的な Intelligence Index タスクを完了するまでの時間は Solar Pro 4 で 8.6 分、Solar Pro 3 では 6.0 分かかります。
料金は入力・出力ともに 100 万トークンあたり 0.30 ドル/1.20 ドルです。これは Intelligence Index のスコアが 45 と 3 ポイント高い MiniMax-M3(自社価格)と同等水準ですが、Intelligence Index で 52 を記録し、推論・最大効率モードで 100 万トークンあたり 0.14 ドル/0.28 ドルの DeepSeek V4 Flash 0731 よりも高価です。キャッシュヒット料金は 100 万トークンあたり 0.06 ドルで、入力トークン価格の 80% オフとなります。
追加モデル詳細:
➤ コンテキストウィンドウ: 384K トークン
➤ 最大出力トーク数: 256K
➤ モダリティ: テキスト入力・出力のみ
➤ 料金: 100 万トークンあたり 0.30 ドル(入力)/ 1.20 ドル(出力)/ 0.06 ドル(キャッシュヒット)
➤ 推論プロバイダー(ローンチ時): Upstage 自社 API、OpenRouter

Solar Pro 4 の最も顕著な改善点は、実世界のエージェントタスクです。GDPval-AA v2 での Elo スコアは Solar Pro 3 の 498 から 1277 に向上しました。このスコアは人間の Elo ベースライン(1000)を上回り、Qwen3.7 Max(1272)や MiMo-V2.5-Pro(1266)をわずかに上回っています。

Solar Pro 4 の AA-Omniscience スコアは -53 から -1 に改善しましたが、これは知識の増加によるものではなく、回答を保留する(アブステイン)割合が増えたことが主な要因です。同モデルが回答を試みた質問の割合は Solar Pro 3 の 92% に対し 41% と大幅に低下しています。一方、ハルシネーション(幻覚・誤答)率は 88% から 24% に改善しましたが、AA-Omniscience Accuracy は 19% でほぼ横ばいとなっています。

Solar Pro 4 は、Artificial Analysis Intelligence Index のタスクで 1 つあたり 43,000 トークンを出力します。これは Solar Pro 3 が 52,000 トークンだったものより約 17% 少ないですが、同レベルのモデルと比較すると依然として冗長な出力となっています。

Solar Pro 4 は、Artificial Analysis の知能指数タスクを 1 つ処理するのに 8.6 分かかります。これは Solar Pro 3 の 6.0 分よりも遅い結果ですが、タスクあたりの出力トークン数は減少しています。

Artificial Analysis Intelligence Index におけるすべての結果:

原文を表示
Korean AI lab Upstage has released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, a significant increase from Solar Pro 3’s 14
Solar Pro 4 is Upstage AI's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026. At 42 on the Intelligence Index it sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43), and shows a 27-point increase over Solar Pro 3. Pricing increases to $0.30/$1.20/$0.06 per 1M input/output/cache hit tokens from Solar Pro 3's $0.15/$0.60/$0.02 via Upstage’s first-party API.
Key results:
➤ Solar Pro 4’s largest improvements on Solar Pro 3 are on agentic and long context work. Terminal-Bench v2.1 improves from 12% to 57%, AA-LCR from 31% to 71%, and τ³-Banking from 9% to 23%. GDPval-AA v2 shows strong progress on real-world agentic tasks, where Solar Pro 3 scored an Elo of 498, well below the human baseline of 1000, Solar Pro 4 scores 1277.
➤ AA-Omniscience improvement from -53 to -1 was from abstaining on more questions. Solar Pro 4 attempts only 41% of questions against 92% for Solar Pro 3, and its hallucination rate is 24%, higher than Command A+ (14%) and MiniMax-M3 (18%), and a vast improvement from Solar Pro 3’s 88%. AA-Omniscience Accuracy remains unchanged at 19%.
➤ Solar Pro 4 is more token efficient than Solar Pro 3, though still verbose for its intelligence level. It uses 43k output tokens per Intelligence Index task, around 17% fewer than Solar Pro 3's 52k.
➤ The intelligence gain comes with a hit to latency. Solar Pro 4 takes 8.6 minutes to complete an average Intelligence Index task, against 6.0 minutes for Solar Pro 3, despite using fewer output tokens per task.
➤ Pricing is $0.30/$1.20 per 1M input/output tokens. This is in line with MiniMax's first-party pricing for MiniMax-M3, which scores 3 points higher at 45, and is more expensive than DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at $0.14/$0.28 and 52 on the Intelligence Index. Cache hits are priced at $0.06 per 1M, an 80% discount on input token price.
Additional model details:
➤ Context window: 384K tokens
➤ Max output tokens: 256K
➤ Modalities: Text input and output only
➤ Pricing: $0.30 / $1.20 / $0.06 per 1M input/output/cache hit tokens
➤ Inference providers at time of launch: Upstage first-party API, OpenRouter

Solar Pro 4's strongest improvement is on real-world agentic work. It scores an Elo of 1277 on GDPval-AA v2, up from 498 for Solar Pro 3. Solar Pro 4 sits above the human Elo baseline of 1000, slightly ahead of Qwen3.7 Max (1272) and MiMo-V2.5-Pro (1266).

Solar Pro 4's AA-Omniscience score improves from -53 to -1, however the improvement comes from abstention rather than knowledge. It attempted only 41% of questions compared to 92% for Solar Pro 3, and while its hallucination rate improves from 88% to 24%, its AA-Omniscience Accuracy is almost unchanged at 19%.

Solar Pro 4 uses 43k output tokens per Artificial Analysis Intelligence Index task, ~17% fewer than Solar Pro 3's 52k. However, it is still verbose compared to other models for its intelligence level.

Solar Pro 4 takes 8.6 minutes per Artificial Analysis Intelligence Index task, up from 6.0 minutes for Solar Pro 3, despite using fewer output tokens per task.

Full results across across the Artificial Analysis Intelligence Index:

News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み