NVIDIA、高性能小型オープンウェイトモデル「Nemotron 3.5 Lightning」を公開
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
NVIDIA は小規模ながら高性能なオープンウェイトモデル「Nemotron 3.5 Lightning」をリリースし、推論速度とエージェント性能で大幅な進歩を達成した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 22:58
AI深層分析
キーポイント
高い知能指数と効率性の両立
Artificial Analysis の知能指数で 24 を記録し、パラメータ数の約 1/4 で同等性能を持つ GPT-OSS-120B に匹敵する成果を示した。
驚異的な推論速度
NVFP4 量子化モデルのテストで秒間約 670 トークンの出力速度を達成し、同規模の競合モデルよりも格段に高速である。
エージェント性能の飛躍的向上
GDPval-AA v2 や Terminal-Bench v2.1 などの評価で前世代比大幅なスコアアップを記録し、高負荷なエージェント展開に適している。
商用利用可能なオープンライセンス
OpenMDW-1.1 ライセンスの下で公開され、実質的な制限なく商業利用が可能となっている。
圧倒的な推論速度
Nemotron 3.5 Lightning は事前リリースの DeepInfra エンドポイントで約 0.5 分のタスク処理時間を達成し、Qwen や Gemma のオープンウェイトモデルを大幅に上回る。
重要な引用
NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters
In pre-release testing of a DeepInfra endpoint serving the final NVFP4 weights, we measured median output speeds of nearly 670 tokens per second
the largest improvements over Nemotron 3 Nano come on agentic evaluations in GDPval-AA v2 (+334 ELO, moving past gpt-oss-120b and Nemotron 3 Super)
Nemotron 3.5 Lightning is highly performant: across the Artificial Analysis Intelligence Index, its Time per Intelligence Index Task is ~0.5 minutes based on output speeds achieved with a pre-release DeepInfra endpoint and the final NVFP4 weights.
編集コメントを表示
編集コメント
NVIDIA は小規模モデルの性能限界をさらに押し上げることに成功しており、エッジデバイスや高スループットな実用システムへの展開が加速する。特に NVFP4 量子化による高速化は、リアルタイム性が求められるアプリケーションにとって決定的な利点となるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
NVIDIA は本日、高性能な小型オープンウェイトモデル「Nemotron 3.5 Lightning」をリリースしました。このモデルは、約 120B のパラメータを持つ gpt-oss-120b に匹敵する性能を持ちながら、総パラメータ数はその約 4 分の 1 です。
Nemotron 3.5 Lightning は、NVIDIA Nemotron 3 Nano 30B A3B の後継モデルです。総パラメータ数は 316 億(31.6B)、アクティブなパラメータ数は 36 億(3.6B)で、Nemotron 3 Nano と同じくハイブリッド Mamba-Transformer アーキテクチャを採用し、小型サイズを維持しています。しかし、知能レベルとエージェントとしての性能には大幅な向上が図られています。
主なポイント:
➤ 知能の飛躍的な向上: Nemotron 3.5 Lightning は Artificial Analysis のインテリジェンス指数で 24 を記録し、Nemotron 3 Nano(15)から 9 ポイントも改善されました。このスコアは OpenAI の gpt-oss-120b(24)と同等であり、約 4 倍の規模を持つ Nemotron 3 Super(26)にも僅差で迫るものです。
➤ 効率性の最適化: Nemotron 3.5 Lightning は、Qwen3.6 35B A3B(スコア 32)や Muse Glimmer(高評価、スコア 35)といった同クラスの最優秀モデルには知能面で劣りますが、異なるフロンティアを目指して設計されています。最終的な NVFP4 重みを実装した DeepInfra エンドポイントでのプレリリーステストでは、出力速度の中央値が秒間約 670 トークンに達することが確認されました。これは、現在市場で提供されている同規模モデルよりもはるかに高速です。
➤ 意味のあるエージェント性能の向上: Nemotron 3 Nano と比較した際、最大の改善が見られたのは GDPval-AA v2 および Terminal-Bench v2.1 におけるエージェント評価です。GDPval-AA v2 では ELO が +334 上昇し、gpt-oss-120b や Nemotron 3 Super を抜いて首位となりました。Terminal-Bench v2.1 でも 7% から 24% と大幅なスコアアップを記録しました。これに高速処理と、制限の少ない OpenMDW-1.1 ライセンスが加わることで、Lightning は高負荷のエージェント展開に適した効率的なワークホースモデルとしての地位を確立しています。
➤ ほぼ損失のない NVFP4 量子化: これまでの Nemotron リリースと同様、本モデルは BF16 重みと併せて NVFP4 形式でも提供されます。Intelligence Index での測定では NVFP4 バリアントがスコア 24 を記録し、高精度な重みと比較しても性能低下は最小限に抑えられています。
主要なモデル詳細:
➤ トークン数 100 万のコンテキストウィンドウを備えた、テキスト特化型の推論モデル
➤ パラメータ数は合計 316 億個(総計)、アクティブに使用されるのは 36 億個
➤ OpenMDW-1.1 ライセンスの下で公開されており、実質的な制限なく商用利用が可能
➤ モデル重みは現在入手可能であり、DeepInfra、Fireworks、FriendliAI、CoreWeave、GMI Cloud、Nebius、Crusoe といったプロバイダーによるサーバーレス推論も利用できます。

Nemotron 3.5 Lightning は極めて高いパフォーマンスを発揮します。Artificial Analysis の Intelligence Index によると、事前リリース版の DeepInfra エンドポイントと最終的な NVFP4 重みを用いて測定した出力速度に基づけば、知能指数タスクあたりの所要時間は約 0.5 分です。
このモデルは、オープンウェイトの競合他社を大きく上回る速度で動作します。具体的には、Qwen3.6 35B A3B(約 3.5 分)、gpt-oss-120b(約 3.4 分)、Gemma 4 31B(約 5.8 分)、そして Qwen3.6 27B(約 7.3 分)を大きく引き離しています。
NVIDIA は CodeRabbit や Harvey といったパートナー企業と協力し、ドメイン固有のタスクでより高いパフォーマンスを発揮できるよう Nemotron 3.5 Lightning をポストトレーニングしました。これにより、手軽な学習プロセスを通じて、ユーザー固有のワークフローに最適な効率と性能を実現することを目指しています。
ただし、全体としての時間効率性においては、依然としてクローズドソースモデルが強い存在感を示しています。Gemini 3.5 Flash-Lite はタスクあたりの所要時間が同程度ながらインテリジェンス指数を 37 に達成しており、GPT-5.6 Luna(最大値)に至っては 2 分未満でスコア 52 を記録しています。

このタスクあたりの所要時間は、極めて高い出力速度と堅牢なトークン効率によって支えられています。Nemotron 3.5 Lightning は、インテリジェンス指数を 9 ポイント向上させたにもかかわらず、Nemotron 3 Nano と同程度の出力トークン数でタスクを完了しています。また、リリース前の DeepInfra エンドポイントでの測定では、1 秒あたり約 670 トークンの出力速度を記録しました。

Nemotron 3.5 Lightning のエージェント機能は、同ファミリーの小型モデルにとって飛躍的な進歩です。GDPval-AA v2 の Elo スコアが 824 に達し、これは Nemotron 3 Super や gpt-oss-120b を上回る性能を示しています。また、Terminal-Bench v2.1 では 24% というスコアを記録しており、Nemotron 3 Nano の 7% と比較すると 3 倍以上の差があり、gpt-oss-120b にほぼ匹敵するレベルです。このサイズと速度を持つモデルにとって、Lightning はエージェントパイプラインに最適な選択肢と言えます。

原文を表示
NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters
Nemotron 3.5 Lightning is the successor to NVIDIA Nemotron 3 Nano 30B A3B, with 31.6B total and 3.6B active parameters. It retains the same hybrid Mamba-Transformer architecture and small size from Nemotron 3 Nano, but makes substantial gains in intelligence and agentic performance.
要点
➤ Major intelligence jump: Nemotron 3.5 Lightning scores 24 on the Artificial Analysis Intelligence Index, a +9 point improvement over Nemotron 3 Nano (15). This puts it in line with OpenAI's gpt-oss-120b (24) and only just behind Nemotron 3 Super (26), a model ~4x its size
➤ Optimized for efficiency: Nemotron 3.5 Lightning sits behind the most intelligent small models in its size class such as Qwen3.6 35B A3B (32) and Muse Glimmer (high, 35) - but it is built for a different point on the frontier. In pre-release testing of a DeepInfra endpoint serving the final NVFP4 weights, we measured median output speeds of nearly 670 tokens per second, much faster than those models are served in the market today
➤ Meaningful agentic gains: the largest improvements over Nemotron 3 Nano come on agentic evaluations in GDPval-AA v2 (+334 ELO, moving past gpt-oss-120b and Nemotron 3 Super) and Terminal-Bench v2.1 (24% vs 7%). Combined with its speed and permissive OpenMDW-1.1 license, this positions Lightning as an efficient workhorse model for high-volume agentic deployments
➤ Near-lossless NVFP4 quantization: as with prior Nemotron releases, the model ships in NVFP4 alongside BF16 weights. We measured the NVFP4 variant at 24 on the Intelligence Index and saw minimal degradation compared to the higher-precision weights
Key model details:
➤ 1 million token context window, text-only reasoning model
➤ 31.6B total and 3.6B active parameters
➤ Released under the OpenMDW-1.1 license, open for commercial use without material restrictions
➤ The model weights are available now along with serverless inference from providers including DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe

Nemotron 3.5 Lightning is highly performant: across the Artificial Analysis Intelligence Index, its Time per Intelligence Index Task is ~0.5 minutes based on output speeds achieved with a pre-release DeepInfra endpoint and the final NVFP4 weights.
This is substantially faster than open weights peers, well ahead of Qwen3.6 35B A3B (~3.5 min), gpt-oss-120b (~3.4 min), Gemma 4 31B (~5.8 min) and Qwen3.6 27B (~7.3 min). NVIDIA has worked with partners including CodeRabbit and Harvey to post-train Nemotron 3.5 Lightning to perform better within domains, aiming to leverage easy training to achieve strong efficiency and performance for user-specific workflows.
However, proprietary models still perform strongly on the overall time-efficiency frontier: Gemini 3.5 Flash-Lite achieves an Intelligence Index of 37 at a similar time per task, and GPT-5.6 Luna (max) scores 52 at under 2 minutes per task.

This time per task is driven by extremely high output speeds along with solid token efficiency. Nemotron 3.5 Lightning used a similar number of output tokens per task to Nemotron 3 Nano while delivering its +9 point Intelligence Index gain, and on the pre-release DeepInfra endpoint we measured output speeds of nearly 670 tokens per second.

Nemotron 3.5 Lightning's agentic capabilities are a step change for the Nemotron family's small models. Its GDPval-AA v2 Elo of 824 surpasses both Nemotron 3 Super and gpt-oss-120b, while its Terminal-Bench v2.1 score of 24% is >3x Nemotron 3 Nano's 7% and almost matches gpt-oss-120b. For a model of this size and speed, this makes Lightning an attractive model for agentic pipelines.

同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み