NVIDIA、Ollama で「Nemotron 3.5 Lightning」を公開しローカル実行可能に
本文の状態
日本語全文を表示中
詳細モードで約5分の本文を読めます。
同じ出来事の情報源
6媒体で確認
vLLM Blog · Artificial Analysis · Ollama Blog · LMSYS Blog · Cline Blog · The Decoder
各社の報じ方を比較 ↓NVIDIA はオラマプラットフォーム上で、ローカル環境で動作するエージェント特化型オープンモデル「Nemotron 3.5 Lightning」の提供を開始した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 02:01
AI深層分析
キーポイント
ハイブリッドアーキテクチャによる高速推論
合計 30B パラメータのうち、1 トークンあたり 3B のアクティブパラメータのみを使用する混合専門家(MoE)構造を採用し、ローカルシステムでの効率的な動作を実現している。
エージェントタスクへの最適化
ファイルの読み込みやツール呼び出し、失敗した処理のリトライなど、多段階にわたるエージェント作業を想定して設計され、開発者が日常的に使用するツールとの連携に焦点を当てている。
大規模コンテキストとデータプライバシー
最大 100 万トークンのコンテキスト長をサポートし、ローカルデバイス上で動作することで、メールやコードベースなどの機密データを外部へ送信することなく処理できる。
柔軟なインフラ対応とカスタマイズ
NVIDIA RTX PC や DGX シリーズ、データセンター、クラウドなどあらゆる環境で動作可能であり、オープンモデルとして特定のタスクに後からトレーニングして適用できる。
ローカル実行による多様なエージェント活用
メールやカレンダー管理、コードレビュー、セキュリティ分析など、高頻度の処理をローカルで完結させ、プライバシーを守りつつ効率的に運用できる。
重要な引用
It's a 30 billion parameter (3B active) open model from NVIDIA built for agents that stay running
Most of these steps don't need a large model. At 3B active parameters per token, it's built for local systems rather than the datacenter
Running locally means your data stays on your device
Nemotron 3.5 Lightning handles the high-volume steps locally, and a larger hosted model picks up the few that need one.
編集コメントを表示
編集コメント
NVIDIA は従来のデータセンター中心の戦略から、エッジやローカル端末での実用性を重視したモデル展開へと舵を切っている。この「Lightning」シリーズは、大規模な計算リソースがなくても高機能な AI エージェントを運用したい開発者にとって重要なマイルストーンとなるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026 年 8 月 11 日
NVIDIA Nemotron 3.5 Lightning が Ollama で利用可能になりました。このモデルは、ユーザーの端末だけで完結して動作します。
NVIDIA が開発したオープンモデルで、パラメータ数は 300 億(1 トークンあたりアクティブなパラメータは 30 億)です。エージェントとして常時稼働し、コンテキストの収集やツールの呼び出し、多段階タスクの処理を担うことを目的に設計されています。
Nemotron 3.5 Lightning は、ファイルの読み込み、ツールの実行、結果のソート、失敗した処理のリトライといったエージェント向けタスクのために作られています。これらのステップの多くは、大規模なモデルを必要としません。1 トークンあたり 30 億パラメータという設計により、データセンターではなくローカルシステムでの利用に最適化されています。ローカルで動作させることで、データを端末内に保持できます。
モデルの特徴
- 作業環境で稼働: 合計 300 億のパラメータを持ちながら、1 トークンあたりアクティブなのは 30 億のみです。ハイブリッドな Mixture-of-Experts アーキテクチャを採用しています。NVIDIA RTX PC や RTX PRO ワークステーション、DGX Spark、DGX Station などのローカル環境で動作します。また、データセンターやクラウド上でも利用可能です。
- エージェントハッチス向けに設計: Nemotron コアリションと共に開発され、コーディング、ツール呼び出し、指示の遵守、複数回の対話など、開発者が日常的に使用するツールに対応してトレーニングされています。
- 1M トークンコンテキスト: 最大 1M トークンのコンテキスト長をサポートし、マルチターンワークフローにおける長いツール履歴を処理できる余地を残します。
- 最適化された推論: マルチトークン予測 (MTP)、DFlash、または DSpark を用いたスペキュラティブ・ディコーディングにより、同等のオープンモデルと比較して最大 4 倍のスループットを実現します。
- カスタマイズ可能: オープンデータセットでトレーニングされたオープンモデルです。特定のタスクにポストトレーニングし、エッジからデータセンターまであらゆる場所で実行できます。
構築できるもの
Nemotron 3.5 Lightning は、以下のワークロードにおいて特に優れています:
- 長期間稼働するパーソナルアシスタント。 メール、カレンダー、プロジェクト管理、予約処理など。ローカルで実行すれば、エージェントはローカルのコンテキストのみを使用し、一切のデータが外部に送信されることはありません。
- コーディング用サブエージェント。 既存のハッチ内でテストの実行、コードベース内の検索、リファクタリングの適用を行います。
- セキュリティ運用。 アラートの補強、インシデントの分類、ログの照会、指標の相関分析、そしてアナリスト向けの構造化された調査結果の作成を担います。
- クラウドに並ぶローカル層。 高ボリュームの処理ステップは Nemotron 3.5 Lightning がローカルで処理し、必要な少数のケースのみを大規模なホストモデルが引き継ぎます。CLI も API も同じです。
- 自分で訓練する専門家。 オープンウェイトとオープンデータセットにより、Nemotron 3.5 Lightning を特定の狭いタスク用にポストトレーニングし、その結果をローカルで実行できます。
はじめに
Ollama をダウンロードした後、お好みのツールで Nemotron 3.5 Lightning を実行してください。
一般的なチャット
ollama run nemotron-3.5-lightning
Claude Code
ollama launch claude --model nemotron-3.5-lightning
OpenClaw
ollama launch openclaw --model nemotron-3.5-lightning
Hermes Agent
ollama launch hermes --model nemotron-3.5-lightning
OpenCode
ollama launch opencode --model nemotron-3.5-lightning
Apple Silicon を搭載するユーザー向けに、Ollama は最良のパフォーマンスを発揮する nemotron-3.5-lightning:30b-mlx モデルを提供しています。
モデルページでは、さらに多くの統合情報 を確認できます。
このパターンは Ollama のクラウド上で動作するモデルでも同様です。エージェントが個別のステップをより大規模なモデルに送信する場合でも、他の設定を変更する必要はありません。
ベンチマーク結果
Nemotron 3.5 Lightning は、同サイズの他社主要オープンモデルと比較して、スループットが 4 倍向上し、タスク完了時間が 30% 短縮されています。また、エージェント作業、コーディング、推論の各タスクにおいて、トップクラスの精度を誇ります。常時稼働するエージェントにとって重要なのはスループットです。1 分あたりの処理ステップ数が増えれば、長時間かかるタスクも素早く完了します。
詳細な結果とテスト設定については、NVIDIA の発表ブログをご覧ください。
原文を表示
August 11, 2026
NVIDIA Nemotron 3.5 Lightning is now available on Ollama, and it runs completely on your own device. It’s a 30 billion parameter (3B active) open model from NVIDIA built for agents that stay running: gathering context, calling tools, and working through multi-step tasks.
Nemotron 3.5 Lightning is made for agentic tasks such as reading a file, calling a tool, sorting a result, and retrying something that failed. Most of these steps don’t need a large model. At 3B active parameters per token, it’s built for local systems rather than the datacenter, and running locally means your data stays on your device.
Model highlights
- Runs where you work: 30B total parameters with only 3B active per token, on a hybrid Mixture-of-Experts architecture. It runs locally on NVIDIA RTX PCs, NVIDIA RTX PRO workstations, NVIDIA DGX Spark and DGX Station, and in the datacenter and cloud.
- Built for agent harnesses: developed with the Nemotron Coalition and trained for the tools developers already use, across coding, tool calling, instruction following and multi-turn work.
- 1M token context: a context length up to 1M, which leaves room for long tool histories across multi-turn workflows.
- Optimized inference: speculative decoding using multi-token prediction (MTP), DFlash or DSpark, offering up to 4x higher throughput than comparable open models.
- Yours to customize: an open model trained on open datasets. Post-train it for a specific task and run the result anywhere, from edge to datacenter.
What you can build
Nemotron 3.5 Lightning excels on the following workloads:
- Long-running personal assistants. Email, calendar, projects and bookings. Running locally, the agent can use local context and none of it is sent elsewhere.
- Coding sub-agents. Running tests, searching the codebase and applying refactors, inside the harnesses you already use.
- Security operations. Enriching alerts, classifying incidents, querying logs, correlating indicators and preparing structured findings for analysts.
- A local tier alongside the cloud. Nemotron 3.5 Lightning handles the high-volume steps locally, and a larger hosted model picks up the few that need one. Same CLI, same API.
- A specialist you train yourself. Open weights and open datasets, so you can post-train Nemotron 3.5 Lightning for one narrow job and run the result locally.
Get started
Download Ollama, then run Nemotron 3.5 Lightning with your tool of choice.
General chat
ollama run nemotron-3.5-lightning
Claude Code
ollama launch claude --model nemotron-3.5-lightning
OpenClaw
ollama launch openclaw --model nemotron-3.5-lightning
Hermes Agent
ollama launch hermes --model nemotron-3.5-lightning
OpenCode
ollama launch opencode --model nemotron-3.5-lightning
For users on Apple silicon, Ollama offers the model with state-of-the-art performance: nemotron-3.5-lightning:30b-mlx.
See more integrations on the model page.
The same pattern works for models running in Ollama’s cloud, so an agent can send an individual step to a larger model without changing anything else.
Benchmarks
Nemotron 3.5 Lightning offers 4x higher throughput and 30% faster task completion time compared to other leading open models of similar size and offers leading accuracy across agentic, coding and reasoning tasks. For agents that stay running, throughput is the number that matters most: more steps per minute means long tasks finish. Full results and test configurations are in NVIDIA’s launch blog.
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
- vLLM BlogvLLM、NVIDIA「Nemotron 3.5 Lightning」の Day-0 サポート開始
- Artificial AnalysisNVIDIA、高性能小型オープンウェイトモデル「Nemotron 3.5 Lightning」を公開
- LMSYS BlogSGLang、NVIDIA Nemotron 3.5 Lightning の Day-0 サポートを追加
- Cline BlogCline に NVIDIA の常時エージェント向けモデル「Nemotron 3.5 Lightning」
- The DecoderNvidia、高速化重視のオープンウェイトモデル「Nemotron 3.5 Lightning」を公開
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み