SGLang、NVIDIA Nemotron 3.5 Lightning の Day-0 サポートを追加
本文の状態
日本語全文を表示中
詳細モードで約13分の本文を読めます。
SGLang が NVIDIA Nemotron 3.5 Lightning モデルの Day-0 サポートを開始し、開発者は高性能な推論スタックを通じてエッジからデータセンターまで常時稼働型エージェントを構築できるようになった。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 23:13
AI深層分析
キーポイント
SGLang の Day-0 サポート開始
SGLang は NVIDIA Nemotron 3.5 Lightning モデルの即時サポートを発表し、OpenAI 互換の高性能推論スタックを通じてエージェントハッチやローカルアシスタントへの接続を可能にした。
ハイブリッド MoE アーキテクチャ
同モデルは NVIDIA Nemotron 3 Ultra から蒸留され、総パラメータ数 30B のうち同時に活性化するのは 3B のみというハイブリッド混合専門家構造を採用している。
多様な展開ターゲットと用途
NVIDIA DGX シリーズや RTX、Jetson など幅広いハードウェアで動作可能であり、金融・セキュリティ・小売などの専門ワークフロー自動化に活用できる。
SGLangのNVIDIA Nemotron 3.5 Lightning対応機能
SGLangは高性能推論ランタイムとして連続バッチ処理、プレフィックスキャッシュ、予測デコーディングをサポートし、OpenAI互換APIを提供する。
BF16チェックポイントを使用した起動コマンド
Dockerコンテナ内でSGLangを起動するには、--mamba-backendにflashinferを指定し、--reasoning-parserと--tool-call-parserでNemotron 3.5およびQwen3 Coderのパーサーを設定する必要がある。
重要な引用
SGLang is excited to announce Day-0 support for NVIDIA Nemotron 3.5 Lightning
Nemotron 3.5 Lightning is built to handle these high-volume tasks.
With SGLang, developers can serve the model through a high-performance, OpenAI-compatible inference stack
"SGLang provides a high-performance serving runtime, continuous batching, prefix caching, speculative decoding, and an OpenAI-compatible API."
編集コメントを表示
編集コメント
NVIDIA の最新モデルと SGLang の連携により、エッジからクラウドまで一貫した高性能なエージェント基盤が実現された。特にパラメータ数の動的制御や多様なハードウェア対応は、実環境での展開コスト削減に寄与する重要な要素である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
SGLang は、ローカルシステム、エッジ、データセンター、クラウドをまたぐ常時稼働型エージェントを支えるために設計されたカスタマイズ可能なオープンモデル「NVIDIA Nemotron 3.5 Lightning」に対する Day-0 サポートを開始することを発表します。
常時稼働型エージェントは、コンテキストの収集、推論、ツールの利用、そして多段階ワークフローにおける適応を担います。複雑なオーケストレーションにはフロンティアモデルが、一方、大量かつ専門的なタスクの効率的な管理には小規模モデルがそれぞれ役割を果たします。
Nemotron 3.5 Lightning は、こうした大量処理タスクに対応するために構築されました。NVIDIA Nemotron 3 Ultra から知識を蒸留し、Nemotron コーリションと共同で開発されたこのモデルは、強力なコーディング能力、ツール呼び出し機能、指示の遵守、多ターン対話機能を備えています。300 億パラメータを持つハイブリッド型混合専門家(MoE)モデルでありながら、実際の推論時には同時に活性化されるのは 30 億パラメータのみです。
Nemotron 3.5 Lightning は、ローカルのパーソナルアシスタントの動力源となり、金融やリスク管理ワークフローの自動化を支援し、サイバーセキュリティ調査を補助し、通信事業の運用最適化に貢献し、小売体験の向上にも寄与します。組織は自社の用語、ポリシー、ツール、ワークフローに合わせて事後学習(ポストトレーニング)を行い、独自の環境へ展開することが可能です。
SGLang を利用すれば、開発者は高性能で OpenAI 互換な推論スタックを通じてこのモデルをサービス化し、エージェント用ハッチやローカルアシスタント、専門的なエンタープライズワークフローへと接続できます。
TL;DR: NVIDIA Nemotron 3.5 Lightning
- アーキテクチャ: ハイブリッド型混合専門家(MoE)アーキテクチャ
- モデルサイズ: 総パラメータ数 30B、アクティブパラメータ数 3B
- コンテキスト長: 最大 100 万トークン
SGLang が NVIDIA Nemotron 3.5 Lightning の Day-0 サポートを追加
- スペキュレティブ・ディコーディング: マルチトークン予測、DFlash、DSpark を活用
- モダリティ: テキスト入力とテキスト出力に対応
- トレーニング: NVIDIA Nemotron 3 Ultra から知識を蒸留し、一般的なエージェントハーンセス向けに訓練済み
- カスタマイズ: オープンドatasets で訓練されたオープンモデル。専門的なワークフローにおけるポストトレーニングもサポート
- デプロイメントターゲット: NVIDIA DGX Spark、DGX Station、RTX PRO、RTX、NVIDIA Jetson、H100、H200、A100、L40S、B200/GB200、B300/GB300
- ローンチ時の利用可能形式: BF16、NVFP4
- はじめに: Hugging Face からモデルウェイトをダウンロードしてください。BF16 と NVFP4
SGLang を使用して Nemotron 3.5 Lightning を実行するには、はじめにクックブック を参照してください。
SGLang によるインストールとクイックスタート
SGLang は、高性能な推論ランタイム、連続バッチ処理、プレフィックスキャッシング、スペキュレティブ・ディコーディング、そして OpenAI と互換性のある API を提供します。以下の基本コマンドは BF16 チェックポイントを使用します。
docker run --rm -it \
--gpus all \
--cap-add SYS_NICE \
--ipc=host \
--network=host \
--entrypoint /bin/bash \
lmsysorg/sglang:dev-nemotron3-5-lightning
sglang serve \
--model-path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
--mamba-backend flashinfer \
--mamba-radix-cache-strategy extra_buffer \
--reasoning-parser nemotron_3 \
--tool-call-parser qwen3_coder
サーバー起動後、OpenAI 互換のクライアントからリクエストを送信してください。
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:30000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Briefly explain: what is SGLang?"},
],
temperature=1.0,
top_p=0.95,
max_tokens=1024,
)
choice = response.choices[0]
print("Reasoning:", choice.message.reasoning_content)
print("Content:", choice.message.content)
スペキュレティブ・ディコーディングによる最適化された推論
Nemotron 3.5 Lightning は、生成されるトークンの質を維持したまま処理速度を向上させるため、マルチトークン予測(MTP)、DFlash、DSpark の 3 つの推測デコーディング手法をサポートしています。
MTP は軽量なモデル統合型予測ヘッドを用いて複数の未来トークンを提案します。DFlash は拡散ベースのドラフトモデルを使って候補ブロック全体を並列生成し、DSpark は信頼度に応じた半自己回帰的なドラフティングを追加することで、速度とトークン採択率のバランスを取ります。これら 3 つの手法を使えば、推論ワークロードに合わせて最適なレイテンシ、スループット、およびデプロイ上のトレードオフを選定できます。
低レイテンシでのサービス提供には H100、H200、DGX Spark 上で DSpark を使用してください。一方、今日時点での最大スループットを追求する場合は、推測デコーディングなしで実行することを推奨します。
MTP を使った Nemotron 3.5 Lightning の実行
Nemotron 3.5 Lightning はマルチトークン予測機能を備えています。SGLang では、モデルに組み込まれた予測ヘッドが未来のトークンをドラフトし、ターゲットモデルがそれらを検証する推測デコーディングパスを通じて MTP を利用できます。
標準的な SGLang の MTP インターフェースでは、EAGLE 推測アルゴリズムを使用します。ハードウェアごとの正確なコマンドについては、cookbook を参照してください。
DFlash を使った Nemotron 3.5 Lightning の実行
DFlash は専用の拡散ドラフトモデルを用いて、ターゲットモデルが並列で検証するトークンの連続ブロックを提案します。SGLang では、DFlash を使用するには互換性のあるドラフトチェックポイントが必要であり、MTP とは別に有効化を行う必要があります。
SGLang の DFlash 実装では、データ並列アテンションには対応しておらず、パイプライン並列サイズは 1 に設定する必要があります。現在のパラメータ参照については、SGLang 推測デコーディングガイドをご覧ください。
DSpark で Nemotron 3.5 Lightning を実行する
DSpark は、自己回帰方式と並列拡散スタイルのドラフティングを組み合わせるハイブリッド型スペキュレーターです。MTP の完全な自己回帰アプローチとも、DFlash の完全な拡散ベースのアプローチとも異なり、DGX Spark 上ではこの 3 つの中で最も高いパフォーマンスを発揮します。
エージェントステップごとの推論制御
Nemotron 3.5 Lightning は、推論機能のオン/オフを切り替え可能です。これにより、ルーターやエージェントハネスは、難易度の高いステップには深い推論を活用し、日常的なタスクには即答で対応させる柔軟な運用が可能になります。
推論機能はデフォルトで有効化されています。推論プロセスを含まない直接の回答を要求する場合は、chat_template_kwargs を介して enable_thinking: false を指定してください。
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Classify this ticket: billing or technical?"}],
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
推論機能を有効にしたリクエストの場合、chat_template_kwargs は省略するか(デフォルトで推論が有効)、または明示的に設定してください。
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Plan the steps to migrate this service."}],
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
このモデルは、推論トークンの予算管理にも対応しています。リクエストごとの推論深度と応答時間を調整するには、enable_thinking と併せて thinking_budget(custom_params を経由)を使用してください。
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Debug why this build is failing."}],
extra_body={
"chat_template_kwargs": {"enable_thinking": True},
"custom_params": {"thinking_budget": 512},
},
)
モデルを組み合わせるシステムでは、推論制御が特に役立ちます。オーケストレーターは、計画やコーディング、曖昧な判断には予算を多めに割り当て、一方、情報抽出、分類、構造化変換などには推論オフモードまたは少量の予算で対応できます。
Nemotron 3.5 Lightning の推論最適化
Nemotron 3.5 Lightning は、重みと予測デコーディングスタックを除き、アーキテクチャ上は Nemotron 3 と同一です。そのため、性能向上の多くはランタイム側で実現されました。SGLang へアップストリームに貢献した主な内容は以下の通りです。
- DSpark の統合。自己回帰と拡散スタイルのドラフティングを融合させたハイブリッド型スペキュレーター「DSpark」を SGLang と Nemotron モデル定義に組み込みました。これにより、MTP や DFlash と併せて、3 つのスペキュレーターから選択できるようになりました。
- 量子化された DSpark ドラフトヘッド。ドラフトヘッドを W4A16 で量子化することで、メモリ使用量とステップごとのレイテンシを削減しつつ、受容率(acceptance rate)は維持しています。これは DGX Spark といったメモリ制約の厳しい環境において特に重要です。
- 同期処理と非同期スケジューリングの実装。ドラフト・検証ループ内のホストとデバイスの同期を排除し、非同期スケジューリングを有効化しました。これにより、現在のバッチが実行されている間に次のバッチの準備が可能になります。
専門 AI 向けに精度と効率性を両立した Nemotron 3.5 Lightning
Nemotron 3.5 Lightning は、ハイブリッド MoE アーキテクチャを採用しています。1 トークンあたり 30B パラメータのうち、実際に動作するのは 3B のみです。これにマルチトークン予測を組み合わせることで計算量を削減し、生成速度を加速させています。
これらの最適化により、同規模のオープンモデルと比較してスループットが最大 4 倍向上しました。その結果、エージェントが専門的なタスクをより迅速に完了できるようになっています。
NVIDIA Nemotron 3 Ultra から知識を凝縮し、主要なエージェントフレームワークで訓練された「Nemotron 3.5 Lightning」は、最先端のエージェント能力をコンパクトで効率的なモデルへと圧縮して提供します。このモデルは、エージェントの生産性、コーディング支援、ツール利用、指示の正確な実行、そして長文脈推論といったあらゆるベンチマークにおいて高い性能を発揮します。
図 1 に示す通り、推論スループットとトークン効率性の向上により、Nemotron 3.5 Lightning は効率性の最前線に位置しています。これにより、常時稼働するエージェントが大量の作業をより高速に完了できるようになります。

図 1: Nemotron 3.5 Lightning は、同等の精度を維持しながらエージェントタスクを最大 30% 高速に完了することで、効率性の最前線をリードしています。
まとめ
NVIDIA Nemotron 3.5 Lightning は、ローカルシステム、エッジデバイス、データセンター、クラウドなどあらゆる環境へ、高速かつカスタマイズ可能なエージェント知能をもたらします。SGLang の Day-0 サポートにより、開発者は高性能で OpenAI と互換性のあるスタックを通じてこのモデルをサービス化できます。また、各エージェントステップごとの推論制御や、DGX Spark などのローカル展開におけるメモリ管理も可能になります。さらに、マルチトークン予測、DFlash、DSpark を活用することで生成速度の加速も実現可能です。
複数のモデル間でワークロードをルーティングするシステムにおいて、Nemotron 3.5 Lightning は、精度、速度、オープン性、そしてデプロイの制御性がすべて重要となる高負荷な専門タスクに対して、開発者にとって魅力的な選択肢を提供します。
始め方
- SGLang を使用して Nemotron 3.5 Lightning を実行するには、cookbook を参照してください。
NVIDIA Nemotron の最新情報については、NVIDIA ニュースを購読するか、NVIDIA AI を LinkedIn、X、YouTube でフォローしてください。また、Discord の Nemotron channel にも参加できます。
謝辞
NVIDIA Nemotron 3.5 Lightning を SGLang に統合するにあたり貢献いただいた皆様に感謝いたします。
NVIDIA: Nirmal Kumar Juluru, Anusha Pant, Amir Klein, Faradawn Yang, Nave Assaf, Ryan Stewart, Alex Steiner, Bita Rouhani, Seong Hee Lee
SGLang チーム
よくある質問(FAQs)
Nemotron 3 Nano と比較して何が新しくなったのですか?
Nemotron 3 Nano は、総パラメータ数 30B、アクティブパラメータ数 3B の効率的なハイブリッド Mamba-Transformer MoE 設計を採用し、1M トークンのコンテキストウィンドウと制御可能な推論を実現しました。Nemotron 3.5 Lightning は、この基盤を踏まえて以下の 3 つの重要な点で進化しています。
- フロンティアモデルの蒸留: Nemotron 3.5 Lightning は Nemotron 3 Ultra から蒸留されており、NVIDIA の最先端エージェントモデルから得た能力を、より小規模なデプロイ環境でも活用できるように転移しました。
- エージェントハッチ最適化: このモデルは、人気のあるエージェントハッチや多段階ワークフロー向けに訓練されています。特にコーディング、ツール利用、指示の遵守、そして専門的なタスク完了に重点を置いています。
- 予測デコーディング: Nemotron 3.5 Lightning は、マルチトークン予測(MTP)、DFlash、DSpark をサポートしており、複数のトークンを並列でドラフトして検証することで生成速度を加速します。
その結果、より短い時間で、より正確にエージェントタスクを完了できるモデルが完成しました。
原文を表示
SGLang is excited to announce Day-0 support for NVIDIA Nemotron 3.5 Lightning, a customizable open model built to power always-on agents across local systems, the edge, the datacenter, and the cloud.
Always-on agents gather context, reason, use tools, and adapt across multi-step workflows. Frontier models handle complex orchestration, while smaller models efficiently manage high-volume, specialized tasks.
Nemotron 3.5 Lightning is built to handle these high-volume tasks. Distilled from NVIDIA Nemotron 3 Ultra and developed with the Nemotron Coalition, it combines strong coding, tool-calling, instruction-following, and multi-turn capabilities in a 30-billion-parameter hybrid mixture-of-experts model that activates only 3 billion parameters at a time.
Nemotron 3.5 Lightning can power local personal assistants, automate financial and risk workflows, support cybersecurity investigations, optimize telecommunications operations, and improve retail experiences. Organizations can post-train and deploy it for their specific terminology, policies, tools, and workflows.
With SGLang, developers can serve the model through a high-performance, OpenAI-compatible inference stack and connect it to agent harnesses, local assistants, and specialized enterprise workflows.
TL;DR: NVIDIA Nemotron 3.5 Lightning
- Architecture: Hybrid mixture-of-experts architecture
- Model size: 30B total parameters, 3B active parameters
- Context length: Up to 1 million tokens
- Speculative Decoding: Multi-token prediction, DFlash, and DSpark
- Modalities: Text input and text output
- Training: Distilled from NVIDIA Nemotron 3 Ultra and trained for popular agent harnesses
- Customization: Open model trained with open datasets, with support for post-training on specialized workflows
- Deployment targets: NVIDIA DGX Spark, DGX Station, RTX PRO, RTX, NVIDIA Jetson, H100, H200, A100, L40S, B200/GB200, and B300/GB300
- Availability at launch: BF16, NVFP4
- Get started:
Download the model weights from Hugging Face: BF16 and NVFP4
- Run Nemotron 3.5 Lightning with SGLang using the getting-started cookbook
Installation and Quick Start with SGLang
SGLang provides a high-performance serving runtime, continuous batching, prefix caching, speculative decoding, and an OpenAI-compatible API. The following baseline command uses the BF16 checkpoint:
docker run --rm -it \
--gpus all \
--cap-add SYS_NICE \
--ipc=host \
--network=host \
--entrypoint /bin/bash \
lmsysorg/sglang:dev-nemotron3-5-lightning
sglang serve \
--model-path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
--mamba-backend flashinfer \
--mamba-radix-cache-strategy extra_buffer \
--reasoning-parser nemotron_3 \
--tool-call-parser qwen3_coder
After the server starts, send a request with any OpenAI-compatible client:
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:30000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Briefly explain: what is SGLang?"},
],
temperature=1.0,
top_p=0.95,
max_tokens=1024,
)
choice = response.choices[0]
print("Reasoning:", choice.message.reasoning_content)
print("Content:", choice.message.content)
Optimized Inference with Speculative Decoding
Nemotron 3.5 Lightning supports three speculative decoding techniques—Multi-Token Prediction (MTP), DFlash, and DSpark—to accelerate token generation while preserving the target model's output quality.
MTP uses lightweight, model-integrated prediction heads to propose several future tokens; DFlash uses a diffusion-based drafter to generate an entire candidate block in parallel; and DSpark adds confidence-aware, semi-autoregressive drafting to balance speed with token-acceptance quality. Together, they let teams choose the best latency, throughput, and deployment trade-off for their inference workload.
For low-latency serving, use DSpark across H100, H200, and DGX Spark; for maximum throughput today, we recommend running without speculative decoding.
Run Nemotron 3.5 Lightning with MTP
Nemotron 3.5 Lightning includes multi-token prediction. SGLang exposes MTP through its speculative-decoding path, where the model's built-in prediction heads draft future tokens and the target model verifies them.
The standard SGLang MTP interface uses the EAGLE speculative algorithm. See the cookbook for the exact per-hardware command.
Run Nemotron 3.5 Lightning with DFlash
DFlash uses a dedicated diffusion draft model to propose a linear block of tokens that the target model verifies in parallel. In SGLang, DFlash requires a compatible draft checkpoint and is enabled separately from MTP.
SGLang's DFlash implementation does not support data-parallel attention and requires pipeline parallel size 1. See the SGLang speculative-decoding guide for the current parameter reference.
Run Nemotron 3.5 Lightning with DSpark
DSpark is a hybrid speculator that combines autoregressive and parallel diffusion-style drafting, sitting between MTP's fully autoregressive approach and DFlash's fully diffusion-based one, and delivers the best performance of the three on DGX Spark.
Control Reasoning for Each Agent Step
Nemotron 3.5 Lightning supports reasoning on or off, allowing a router or agent harness to use deeper reasoning for difficult steps and direct answers for routine work.
Reasoning is enabled by default. To request a direct answer without a reasoning trace, pass enable_thinking: false via chat_template_kwargs:
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Classify this ticket: billing or technical?"}],
extra_body={"chat_template_kwargs": {"enable_thinking": False}},
)
For reasoning-enabled requests, either omit chat_template_kwargs (reasoning is on by default) or set it explicitly:
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Plan the steps to migrate this service."}],
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
The model also supports a reasoning-token budget. Use thinking_budget (via custom_params) alongside enable_thinking to change the reasoning depth and response time per request:
response = client.chat.completions.create(
model="nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16",
messages=[{"role": "user", "content": "Debug why this build is failing."}],
extra_body={
"chat_template_kwargs": {"enable_thinking": True},
"custom_params": {"thinking_budget": 512},
},
)
Reasoning control is particularly useful in systems of models: the orchestrator can allocate a larger budget to planning, coding, and ambiguous decisions, while using reasoning-off mode or a small budget for extraction, classification, and structured transformations.
Optimizing inference for Nemotron 3.5 Lightning
Nemotron 3.5 Lightning is architecturally identical to Nemotron 3 apart from the weights and the speculative decoding stack, so most of the performance work landed in the runtimes themselves. Here's what we contributed upstream to SGLang:
- DSpark integration. We wired DSpark—a hybrid speculator that blends autoregressive and diffusion-style drafting—into SGLang and the Nemotron model definition, giving you three speculators to choose from alongside MTP and DFlash.
- Quantized DSpark draft head. Quantizing the draft head to W4A16 cuts its memory footprint and per-step latency without hurting acceptance rate, which matters most on memory-constrained parts like DGX Spark.
- Removal of syncs and async scheduling. We eliminated host-device syncs in the draft-and-verify loop and enabled async scheduling, so the next batch is prepared while the current one is still executing.
Nemotron 3.5 Lightning offers Leading Accuracy and Efficiency for Specialized AI
Nemotron 3.5 Lightning combines a hybrid MoE architecture—with only 3B of its 30B parameters active per token—with multi-token prediction to reduce computation and accelerate generation. These optimizations deliver up to 4x higher throughput than similarly sized open models, helping agents complete specialized tasks faster.
Distilled from NVIDIA Nemotron 3 Ultra and trained across popular agent harnesses, Nemotron 3.5 Lightning transfers frontier-level agentic capabilities into a compact, efficient model. It excels across benchmarks for agent productivity, coding, tool use, instruction following, and long-context reasoning.
As shown in Figure 1, higher inference throughput and token efficiency places Nemotron 3.5 Lightning on the efficiency frontier, helping always-on agents finish high-volume work faster.

Figure 1: Nemotron 3.5 Lightning leads the efficiency frontier by completing agentic tasks up to 30% faster at comparable accuracies.
まとめ
NVIDIA Nemotron 3.5 Lightning brings fast, customizable agentic intelligence to local systems, the edge, the datacenter, and the cloud. With SGLang Day-0 support, developers can serve the model through a high-performance, OpenAI-compatible stack; control reasoning per agent step; manage memory for local deployments such as DGX Spark; and accelerate generation with multi-token prediction, DFlash, or DSpark.
For systems that route work across multiple models, Nemotron 3.5 Lightning gives developers a compelling option for high-volume specialized tasks where accuracy, speed, openness, and deployment control all matter.
Get Started
- Download the model weights from Hugging Face: BF16 and NVFP4
- Run Nemotron 3.5 Lightning with SGLang using the cookbook
*Stay up to date on NVIDIA Nemotron by subscribing to NVIDIA news and following NVIDIA AI on LinkedIn, X, YouTube, and the Nemotron channel on Discord.*
Acknowledgement
Thanks to everyone who contributed to bringing NVIDIA Nemotron 3.5 Lightning to SGLang.
NVIDIA: Nirmal Kumar Juluru, Anusha Pant, Amir Klein, Faradawn Yang, Nave Assaf, Ryan Stewart, Alex Steiner, Bita Rouhani, Seong Hee Lee
SGLang Team
FAQs
What is new compared with the Nemotron 3 Nano?
Nemotron 3 Nano established an efficient hybrid Mamba-Transformer MoE design with 30B total parameters, 3B active parameters, a 1M-token context window, and controllable reasoning. Nemotron 3.5 Lightning builds on that foundation in three important ways:
- Frontier-model distillation: Nemotron 3.5 Lightning is distilled from Nemotron 3 Ultra, transferring capabilities from NVIDIA's frontier agentic model into a much smaller deployment footprint.
- Agent-harness optimization: Nemotron 3.5 Lightning is trained for popular agent harnesses and multi-turn workflows, with an emphasis on coding, tool use, instruction following, and specialized task completion.
- Speculative decoding: Nemotron 3.5 Lightning supports multi-token prediction (MTP), DFlash, and DSpark to accelerate generation by drafting and verifying multiple tokens in parallel.
The result is a model designed to complete more agent tasks more accurately in less time.
同じ出来事を5媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み