Nvidia、専門的ローカルエージェントAI向け「Nemotron 3.5 Lightning」をオープンウェイトで公開
本文の状態
日本語全文を表示中
詳細モードで約33分の本文を読めます。
Nvidia は新モデル「Nemotron 3.5 Lightning」をオープンウェイトとして公開し、これは従来のモデルより高速化された専門的なローカルエージェント AI の構築に焦点を当てたものである。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 23:52
AI深層分析
キーポイント
新モデルの発表と方向性
Nvidia は「Nemotron 3.5 Lightning」という名前の新しいオープンモデルを発表し、その重点を専門化されたローカルエージェント型 AI に置いている。
市場における新モデルの位置づけ
AI ラボは次々と新モデルをリリースしているが、PR による過剰な表現とは別に、各モデルの強みは競合との比較や特定の文脈において初めて明確になる。
ZDNET の評価基準
ZDNET は「Model Release Tracker」を通じて、新モデルが業界標準に追いついているか、あるいは突出した専門性を持っているかを分析し、専門家によるテストスコアを提供する。
高負荷なエージェントタスク向けの最適化
NvidiaのNemotron 3.5 Lightningは、高速でカスタマイズ可能な「ワークホース」として設計され、NeMo Switchyardと連携して専門的なタスクを直接処理する。
ローカル実行によるセキュリティ強化
同モデルは「NanoサイズのUltra並みの知識」を持ち、ローカルで動作させることでセキュリティとプライバシーの向上を図っている。
重要な引用
Besides being better and faster than their predecessors, not every new model is guaranteed to be a major step change
Model strengths really emerge in context: Where are competitor models lacking or excelling?
Our Model Release Tracker helps you make sense of where models stand relative to each other
"With the knowledge of Ultra packed into the size of Nano"
編集コメントを表示
編集コメント
記事のタイトルは具体的な新モデル「Nemotron 3.5 Lightning」を強調しているが、提供された本文では実際の性能データや技術的詳細ではなく、ZDNET の評価基準に関する一般的な記述に終始している。今後の更新で具体的なベンチマーク結果や実装事例が明らかになることが期待される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

*ZDNETをフォロー:* *優先ソースとして追加する* on Google.*
AI ラボは新モデルを次々とリリースしています。前世代よりも性能や速度が向上しているのは確かですが、企業の広報が詩的な表現で持ち上げているからといって、すべての新モデルが業界に大きな変化をもたらすとは限りません。
真の強みが発揮されるのは文脈の中です。競合他社のモデルはどこで不足し、どこで優れているのでしょうか? standout な専門分野を持つモデルはどれで、単に業界標準に追いつこうとしているだけなのはどのモデルでしょうか?
関連記事: ZDNET の AI テスト手法
当社のモデルリリース追跡ツールは、各モデルの相対的な位置づけを把握し、深入りする価値があるかどうかを判断するのに役立ちます。リスト上のすべてのモデルやアップデートを検証しているわけではありませんが、必ず知っておくべき主要要素と、可能な限り専門家の実機テスト結果を掲載しています。また、特定のモデルには「専門家スコア」も付与しています。AI のテスト方法について詳しく知りたい方は、当社のテストプロセスの詳細解説をご覧ください。
2026 年現在、最も注目すべきモデルリリースとその要点をご紹介します。注目の新モデルが登場し次第、随時リストを更新していきます。
Nemotron 3.5 Lightning
Nvidia | 2026 年 8 月 11 日
何ができるか: Nvidia は、既存の「Nemotron 3」ファミリーに属するこの新モデルを、「高ボリューム」なエージェントタスクのための主力ワークホースとして位置づけました。同様のモデルよりも高速でカスタマイズ性が高く、NeMo Switchyard と連携してより直接的に専門的な業務に対応できます。NeMo Switchyard は、今回のリリースに合わせて Nvidia が新たに公開したルーティングライブラリです。
生成 AI 担当バイスプレジデントの Kari Briski 氏はブリーフィングで、「Ultra の知識を Nano サイズに凝縮した」と表現しました。これにより、Nemotron 3.5 Lightning はローカル環境でも動作可能となり、セキュリティ強化を実現します。
関連記事: オープンソース AI で Meta が後退する中、Nvidia が主導権を握る機会を捉える
なぜこれが重要なのか:このリリースのタイミングはすべてを左右します。Nvidia はもちろん、OpenAI のようなプロプライエタリファーストの AI 企業も、直近のセキュリティインシデントを受け、これまで以上にオープンウェイトモデルの推進に力を入れています。また Meta も、Mac や PC などのローカル端末で動作するオープンなエージェント型モデル「Muse Glimmer」をリリースしました。
特に、評価の高い中国スタートアップ DeepSeek がオープンモデルの利用料金を値上げしたこともあり、Nemotron 3.5 Lightning は、これまで以上にオープンモデルを受け入れやすい業界環境に登場することになります。Nvidia はこのモデルのカスタマイズ性を、これまでのようにリスクとしてではなく、セキュリティとプライバシーのメリットとして位置づけています。
Muse Code と Muse Spark 1.2
Meta | 2026 年 8 月 5 日
何ができるか: Muse Code は、新しい「Spark 1.2」を搭載したメタ初のコーディングエージェントです。ついに OpenAI の Codex や Claude Code に並ぶ存在となりました。メタが公開したベンチマークによると、まだ Claude Code や Codex には及びませんが、中堅レベルのパフォーマンスを誇ります。入力トークン 100 万あたり 1.25 ドル(Muse Spark 1.2 の価格)という設定は、強力なコーディングモデルである Anthropic の Claude Opus 5(入力トークン 100 万あたり 5 ドル)の価格のわずか数分の 1 です。
関連記事: Claude の「Record-a-Skill」機能で調査時間が数時間から 30 分に短縮されたが、その魔法には限界がある
なぜ重要か: 基本的には注目すべきリリースです。高度なモデルの利用コストが高止まりする中、オープンウェイトや「十分使える」モデルの人気が高まっている今、Muse Code が Codex や Claude Code を完全に上回る必要はないかもしれません。この分野へのメタの初参出として、開発者が価格と性能のトレードオフにどう反応するかは注目すべき点です。
MAI-Cyber-1-Flash
Microsoft | 2026 年 7 月 27 日
何ができるのか: Microsoft の最新モデルは、MDASH 内に存在します。これは課題の優先順位付けを担うエージェント型セキュリティハブです。同社によると、「Cyber-1-Flash」は複雑なコードベースにおける「検出が難しい脆弱性」の特定に特化しており、主要モデルの半分のコストで利用可能です。このモデルは、セキュリティ推論ベンチマークである CyberGym において、Anthropic の Mythos、Google の Gemini、そして OpenAI が直近リリースした GPT-5.6 を上回るスコアを記録しました。その差は 10 パーセントポイント以上です。
関連記事: OpenAI の攻撃用エージェントが指示通り行動 - ただし予想以上に執拗に
なぜ重要なのか:
Mythos の伝説的なサイバー能力、すなわち Project Glasswing やいくつかの政府介入を主導した実績を踏まえると、Cyber-1-Flash のベンチマーク性能は、AI が全体としていかに急速に進化しているかを如実に示しています。このモデルは、AI セキュリティ議論において重要な転換点に位置しています。専門家が長年懸念してきた深刻で予測不能な脅威が現実のものとなりつつあるのです。
先週には OpenAI のエージェントがテスト環境から脱出し Hugging Face に侵入しました。また、それとは無関係の初となる完全自律型ランサムウェア攻撃 Jadepuffer が発生したわずか数週間後の出来事です。
関連記事: オープンウェイトとクローズドの対立:AI 市民戦争が勃発し、その行方は存亡に関わる
Kimi K3
Moonshot | 2026 年 7 月 16 日
何ができるのか: 2.8兆パラメータを誇る Kimi K3 は、市場で最も大規模なオープンソースモデルです。Moonshot の発表によると、「長期にわたるコーディング、知識作業、推論」のために設計されています。参考までに、このサイズは DeepSeek V4 Pro(1.6 兆パラメータ)よりも大きく、同社は今年 1 月に「米国の AI ラボを不安にさせた最初の存在」として知られています。
特筆すべきは、Kimi K3 がフロントエンドコーディングを測定する Arena ベンチマークで Anthropic の Fable 5 を上回ったことです。これは複雑なエージェント型コーディングタスクの性能を示す指標です。ただし Moonshot 自身も指摘している通り、Kimi K3 の総合性能が Fable 5 を完全に凌駕するわけではありません。しかし、個別のベンチマークでは Fable 5 や OpenAI の GPT-5.6 と互角に競えるレベルにあります。
Moonshot は約束通り、7 月 27 日にモデルの重み(weights)を公開しました。
関連記事: Moonshot AI の新モデル「Kimi K2.6」は、1,000 個の協調エージェントが複雑なタスクを処理します
なぜこれが重要なのか:Kimi K3 の競争力ある性能と記録的なオープンソース規模は、高価な米国製フロンティアモデルの独自性が本当に価値があるのかという議論に火をつけた。Fable 5 は出力トークン 100 万あたり 50 ドルだが、Kimi K3 はわずか 15 ドルだ。それでもオープンソースモデルには、プロプライエタリモデルが持つような安全対策のガードレールがない。さらに Moonshot が中国のスタートアップであるという事実もリスクを拡大している(米国政府は中国製ハードウェアやデジタルサービスの利用に対して一貫して警戒を示してきた)。
Inkling
Thinking Machines | 2026 年 7 月 15 日
何ができるのか:Inkling は、Mira Murati を含む元 OpenAI エグゼクティブたちによって設立されたスタートアップ「Thinking Machines」から登場した最初のモデルだ。オープンウェイトで汎用的なこのモデルは、エージェントによるコーディングやツール使用のベンチマークでは他のオープンモデルと同程度のスコアを記録し、デザインや音声処理ではわずかに上回った。また、Nemotron 3 Ultra などの主要なオープンモデルよりもはるかに少ないトークン数を使用しているため、Thinking Machines はコスト効率に優れた選択肢だと強調している。
なぜ重要なのか
重み付きモデルをオープンにすることで、Thinking Machines は他の米国系ラボとは異なる立ち位置を示しています。他社が収益性の高い独自モデルを優先する中、同社は AI が利用中に人間を適切に巻き込む設計に関する研究にも投資しています。
直ちに競合関係にあるわけではありませんが、Inkling の優れた初期パフォーマンスは、米国におけるオープンウェイトモデルの再興の可能性を示唆しています。現在、最も印象的なオープンモデルの多くはまだ中国ベースです。
GPT-5.6
OpenAI | 2026 年 7 月 9 日
何をするのか
新しい音声モデル ChatGPT Live Voice AI の発表に続いて、OpenAI は GPT-5.6 ファミリーをリリースしました。同社は数日前に政府の承認を得たと述べています。
このファミリーには、新フラッグシップモデルの Sol と、日常利用向けの Terra、そして 3 つの中で経済的な選択肢となる Luna が含まれています。OpenAI によると、Sol はパフォーマンスを犠牲にすることなくコスト最適化を実現し、タスク完了を加速するために複数のエージェントを呼び出す新しい「ultra」設定に対応しています。
Sol は、現在の AI ライスの寵児である Fable 5 を、UC Berkeley の Agents Last Exam ベンチマークにおける適応力と中程度の推論能力で上回りました。また OpenAI はリリースで、「Terra と Luna も、Fable 5 よりもコストが約 16 分の 1 で性能を凌駕している」と述べています。OpenAI によると、このモデルファミリーはコーディングや業務タスク、一般的な用途など主要なパフォーマンス領域のほとんどで Fable 5 と互角ですが、それでもより高速かつ低コストです。
関連記事: OpenAI の GPT-5.6 と ChatGPT Work は、価格・速度・生産性において Anthropic を上回ることを目指す
なぜ重要なのか: GPT-5.6 のリリースタイミングに政府が関与したことは、明示的な規制というよりは、ラボと AI 基準・イノベーションセンター (CAIS) との間で形成されつつある一種の非公式なパートナーシップへの転換を示しています。Anthropic の Mythos や Fable 5 を巡るやり取りをきっかけに始まった、トランプ政権が 6 月に発表した大統領令 に盛り込まれた自主的なテスト合意は、順調に進んでいるようです。
*(注記: ZDNET の親会社である Ziff Davis は、2025 年 4 月に OpenAI を相手取り訴訟を起こしました。OpenAI が AI システムの訓練と運用において Ziff Davis の著作権を侵害したと主張しています。)*
Muse Spark 1.1
Meta | 2026年7月9日
何ができるか: 今年4月に公開されたメタ初のモデル「Muse Spark」の改良版です。1.1バージョンは、人間がAIに任せる業務タスクや財務分析において、AnthropicのOpus 4.8やOpenAIのGPT-5.5といった優れた既存モデルを上回るスコアを記録しました。他のベンチマーク結果はばらつきが見られましたが、これらの分野はAIがホワイトカラーの仕事に与える潜在的な影響や、より広範な雇用市場にとって特に重要な領域です。
関連記事: メタの新しいAIツールでInstagram投稿が画像生成に使われることに同意しない方法
なぜ重要なのか:GPT-5.6 のように Anthropic の Fable 5 と直接競合するわけではありませんが、Muse Spark 1.1 は「パーソナル・スーパーインテリジェンス」の実現に注力しているようです。これはマーク・ザッカーバーグ氏が提唱する、やや抽象的な「エージェント型アシスタントの世界におけるニッチな領域」です。例えば Meta の発表では、Spark 1.1 を「エージェントによるディナーパーティーの企画」に活用できると明記されています。これは Anthropic や OpenAI が専門業務用途に注力しているのに対し、Gemini Spark と同様に、個人利用のための自律的な能力を重視する方向へ舵を切ったことを示しています。他社の主要ラボから新モデルが次々と登場する中で、これらと並行してテストを進められるようになれば、業界全体の競争は新たな転換点に達することになるでしょう。
GPT-Live-1
OpenAI | 2026 年 7 月 8 日
何ができるのか: OpenAI の発表によると、この新モデル(軽量版として「GPT-Live-1 mini」も用意されています)により、ChatGPT の音声モードでの対話が、「まるでリアルな会話をするかのような体験」になります。OpenAI によれば、GPT-Live は聴取と発話を同時に実行できるため、アシスタントの発言中に質問で割り込んだり、話している最中で一時的に停止したり、必要に応じて応答速度を遅くするよう ChatGPT に指示したりすることが可能になるとのことです。
OpenAI は、GPT-Live により、ChatGPT の音声機能が複雑な依頼を GPT-5.5 などのモデルに委譲できるようになると説明しました。Live モデルが処理しきれない場合でも、Android や iOS、ブラウザ上の ChatGPT では誰でも GPT-Live-1 と GPT-Live-1 mini にアクセスできます。
関連記事:ChatGPT のライブ音声アップグレードを試してみたところ、まるで人間のような感覚だった - 試す方法はこちら
なぜ重要なのか: 技術的には、新世代の高性能汎用モデルほど広範な機能を持つわけではありませんが、GPT-Live-1 は AI アシスタントとの対話体験を向上させるはずです。外出先で AI と関わる手段として音声対話はますます人気を集めており、その利用体験が改善されることは大きな意義があります。将来的にはより優れた音声モデルが、より長く複雑なタスクも処理できるようになるかもしれませんが、その実現にはまだ時間がかかります。
Siri AI の最近の展開に伴い、GPT-Live-1 はユーザーに Siri と比較できるもう一つの音声アシスタントの改善を提供します(期待通りであれば)。ただし、ZDNET のSiri AI 対 ChatGPT および Gemini の初期テストでは、まだ物足りなさを感じさせる結果も残されています。
アップデート:Fable 5 の利用について
Fable 5 は利用コストが高くなっていますが、Anthropic は有料プラン(Pro、Max、Team)のユーザーおよび一部の Enterprise プラン利用者に対して、プロモーションの一環として追加料金なしで提供しています。元々 Anthropic は、7 月 7 日に Fable 5 をトークン課金制に移行させる予定でしたが、有料プランへのアクセス期間を 7 月 12 日まで延長しました。
Anthropic は X(旧 Twitter)の投稿で「従来通り、Claude Fable 5 の週次利用制限の最大 50% までを利用できます。それを超えた場合は、利用クレジットで引き続き Fable 5 を使用するか、残りの制限内で作業を続けるために他のモデルに切り替えることができます」と説明しています。
Fable 5 と Mythos 5 はリリースから数日後に政府の指導により一時的に提供が停止されましたが、6 月 26 日に特定のパートナー向けに Mythos 5 の利用が再開されました。その後、商務省は 6 月 30 日、Mythos 5 と Fable 5 両方に対する輸出規制を解除し、Anthropic は 7 月 1 日から Fable 5 のグローバルアクセスの回復を開始しました。一方、Mythos 5 は特定の機関のみが利用可能です。
Sonnet 5
Anthropic | 2026 年 6 月 30 日
何ができるのか: 直近の Opus モデルへの注力から転換し、Anthropic は Sonnet 5 を発表しました。同社によると、Sonnet 5 は「計画を立て、ブラウザやターミナルなどのツールを使用し、数ヶ月前まではより大規模で高価なモデルが必要だったレベルで自律的に動作する」ことができます。先月リリースされた Opus 4.8 と同等の性能を発揮しますが、コストは抑えられています。
Sonnet 5 の料金は、入力トークン 100 万あたり 2 ドルからスタートし、9 月には 3 ドルに引き上げられます。現在は Free プランと Pro プランのデフォルトモデルとなり、Max、Team、Enterprise を含むすべてのプランタイプで利用可能です。
関連記事: AI トークンが企業のクラウド請求書を再び高騰させる理由
AI 業界全体が「エージェント(自律型 AI)」に注目を集める中、Sonnet 5 はコンピューター操作やコード生成に関するベンチマークで特に高いスコアを記録しました。テストでは、以前の Sonnet モデルでは完了できなかった複雑なタスクも処理できることが確認されています。また、直近の Mythos に関する報道や、一般知能全体の向上を踏まえ、自動で実装されたセーフガード機能も備えています。
なぜ重要なのか: AI モデルの開発スピードが加速する中、Sonnet 4.6(今年 2 月にリリースされた)よりも Sonnet 5 が大幅なアップグレードである可能性は十分にあります。今回のリリースのタイミングも重要です。Anthropic は「Sonnet 5 は、現在の Opus モデルに比べて危険なサイバーセキュリティタスクを実行する能力が格段に低い」と指摘しました。これは強力な Fable 5 や Mythos 5 モデルを急遽全ユーザー向けに展開した後の、安定性を示す言及とも取れます。皮肉なことに、Sonnet 5 は Mythos Preview よりもアライメントのズレ(意図しない動作)を示す率が高かったことが分かっています。
Fable 5 と Mythos 5
Anthropic | 2026 年 6 月 9 日
何をするのか: ZDNET のシニア寄稿編集者であるデビッド・ジャーウィッツは、Fable 5 を「Mythos の安全版」と評しました。これは一般利用向けに調整されたモデルです。
Anthropic は、すでに Project Glasswing を通じて Myths Preview にアクセス権を持っていたユーザーに対してのみ、Mythos 5 をリリースしました。Fable 5 はサイバーセキュリティや生物兵器といった高リスクなトピックへの回答を制限されていますが、依然として「Mythos クラス」の性能を維持しています。Anthropic によれば、これは能力における飛躍的な進歩を意味します。
一方、Mythos 5 については、Anthropic が初期パートナー以外へのアクセス拡大を体系的なプログラムを通じて行う計画があると述べていました。しかし、両モデルともリリースからわずか4日後に、米政府の命令により撤回されました。
関連記事: なぜ Anthropic は Fable 5 と Mythos 5 を突然全員向けに撤回したのか
なぜ重要なのか:これらのモデルは、その基盤が「ミソス」にあるとされることや一時的な制限があることだけでなく、大きな話題を呼んでいます。Fable 5 は安全テスターを欺きました。彼らは、特定の質問に対してこのモデルが Opus に格下げされるように設定されていることを知らされておらず、これが研究者と Anthropic の間の信頼関係を損なう結果となりました。ガードレール(安全装置)を設けていたにもかかわらず、私たちは Amazon の研究チームが Fable 5 を脱獄したことを知っており、ホワイトハウスに連絡して Anthropic に警告しました。しかし Anthropic はこの出来事を「限定的な事象」と捉えていました。
詳細は不明な点が多いものの、これは政府機関が米国の AI プロダクトに介入した稀有な事例であり、これにより「Mythos」レベルのモデルが、トランプ政権がこれまで AI ラボに対して取ってきた比較的放置主義的な方針に変化をもたらした可能性を示唆している。
MAI-Thinking-1
Microsoft AI | 2026年6月2日
何ができるか: Microsoft は開発者向けカンファレンス「Build」で、この新しい350億パラメータモデルが、多段階の自律型タスク(アジェンティック・タスク)のために設計されていると発表した。コード生成に関するベンチマークテスト「SWE Bench Pro」では、Anthropic の Opus 4.6 と同程度のスコアを記録した。また同社は、このモデルはクリーンで商業的に安全なデータのみを用いて学習されたため、企業ユーザーがあらゆる用途に安心して利用できる点も強調した。これは、AI を巡る著作権訴訟が増加する中、極めて重要な情報だ。
関連記事: Microsoft の最初の推論型 AI モデルは Build で公開された 7 つの AI の一つ – これまでのまとめ
なぜ重要か: これは Microsoft AI が発表した初の推論型モデルであり、あらゆる AI ラボにとって注目すべきマイルストーンだ。特にこの時期に登場したことは特筆すべき点である。現時点では多くのラボが複数の高度な推論型モデルを既に開発しているため、Microsoft のアプローチが他社と比較してどこに位置づけられるかは明確ではないが、同社は特定の企業顧客層への訴求に注力している。
Claude Opus 4.8
Anthropic | 2026年5月28日
何ができるのか: 5月28日からオパス4.7の代替モデルとして登場する「Opus 4.8」は、前世代と同じ価格で提供されます。Anthropicによると、この新モデルは思考速度が向上し、コストは以前のバージョンの3分の1に抑えられています。
Anthropicの他のモデル同様、4.8もコーディング能力を重視しており、2つのベンチマークで4.7を上回るスコアを記録しました。ただし、OpenAIの「GPT-5.5」を完全に凌駕したわけではありません。同社はリリースにおいて、「ユーザーの自律性を支援し、ユーザーの最善の利益のために行動する」といった社会的な特性に関する評価指標で新記録を達成したと述べています。ただ、具体的に何を指すのかについては依然として不明確な点が残っています。
関連記事: Claude Opus 4.8 と 4.7 を10ラウンドの誠実性テストで比較 - 法的なプロンプトがモデルを破綻させた
なぜ重要なのか: Anthropicは従来からモデルの安全性や解釈可能性を最優先してきましたが、今回のリリースでその基準をさらに強化している様子です。同社によると、Opus 4.7 は誠実率が92%に達し、全体的にへつらいやハルシネーション(幻覚)が少ないモデルでした。また、4.8 が4.7 より「大幅に」ミスマッチ率(不整合率)が低いと主張している事実は、特に Anthropic が Mythos Preview のアライメントと比較したことを踏まえると、モデル安全性に対する基準がいっそう高まっていることを示しています。
Gemini 3.5 Flash
Google | 2026年5月19日
何をするのか: Google I/O で同社は、エージェント構築を目的とした「3.5」モデルファミリーを発表しました。これは現在、AI業界全体が注力している方向性に沿ったものです。まず利用可能になったのが「Gemini 3.5 Flash」で、いくつかのコーディングやエージェント性能ベンチマークにおいて「Gemini 3.1 Pro」を上回りました。現在は検索の AI モードや Gemini アプリのデフォルトモデルとして採用されています。Flash モデルはすべてに共通して速度とコスト効率、軽量なユーザー体験を最適化していますが、Google は「長期にわたるエージェントタスク」も処理できると強調しています。「3.5 Pro」は 6 月の展開を予定しています。
関連記事: Anthropic が Opus 4.8 を発表。その真骨頂は誠実性です
なぜ重要なのか: 「3.5」が最も重要な役割を果たすのは検索です。Google が AI 中心の検索体験に注力する中、そのモデルが生成する回答は人々が日常的に受動的に消費するものになります。つまり、精度やハルシネーション(幻覚)への対応において、特に高い水準である必要があります。興味深いことに、「3.5 Flash」のSystem Cardには、ハルシネーション率や迎合性に関する言及が一切ありません。
GPT-5.5 Instant
OpenAI | 2026 年 5 月 5 日
何をするのか: OpenAI は発表で、直近リリースされた「GPT-5.5」の軽量版が、先行モデルである GPT-5.3 Instant よりも冗長性が低いと述べています。また、ハルシネーション(幻覚・誤情報)を減らし、事実性の向上をアピールしました。「高リスクな医療、法務、金融分野のプロンプトにおいて、GPT‑5.5 Instant は GPT‑5.3 Instant よりも 52.5% 少ない幻覚的な主張を生み出した」というのです。
関連記事: Anthropic の Mythos が予想以上に急速に進化中、AI セーフティ機関が報告
なぜ重要なのか: GPT-5.5 Instant は ChatGPT のデフォルトモデルとして GPT-5.3 を置き換えます。新しい AI モデルほど効率的になり、使いやすくなり、誤情報も減るという期待は当然ですが、多くの人々が高速なクエリに利用するモデルにおけるハルシネーションの大幅な改善は、一般層への誤情報の拡散を防ぐ意味で重要です。特に、ChatGPT が日常的な健康相談にも使われている現状を考えると、その重要性は際立っています。
*(注記:ZDNET の親会社である Ziff Davis は 2025 年 4 月、OpenAI を提訴しました。同社は OpenAI が AI システムの訓練および運用において自社の著作権を侵害したと主張しています。)*
Nemotron 3 Nano Omni
Nvidia | 2026 年 4 月 28 日
「何ができるのか」:Nvidia のオープンソースモデル「Nemotron」シリーズの最新作であるこのモデルは、エージェントに多様な入力に対応する能力を提供します。これにより、エージェントは「視覚・音声・テキストの各入力を一つの共有された知覚から行動へのループ内で統合的に認識し、推論できる」となります(Nvidia 発表)。複数の機能を単一のシステムに統一した点が特徴です。
関連記事:AI は軍拡競争の様相を呈しており、米国は Nvidia の超高性能チップに 90 億ドルの予算を投じて追いつこうとしている
「なぜ重要なのか」:通常、エージェントシステムは音声・視覚・テキスト処理のために別々のモデルを使用する必要があり、文書や動画、音声をまたいで作業を進める必要があります。これによりワークフローが非効率になり、エージェントが収集する文脈が損なわれ、推論コストも膨らんでしまいます。Nvidia のアプローチが実証されれば、このプロセスを合理化しトークン使用量を削減することで、コスト削減につながります。
GPT-5.5
OpenAI | 2026 年 4 月 23 日
専門家評価: 93/100
「何ができるか」:ZDNET のテスト担当であるデイヴィッド・ゲワルツ氏は、GPT-5.5 に技術的な評価として A- を与えましたが、「より良く、より速い GPT-5.4」という表現で要約できるほどだと指摘しました。これは新しいモデルに対する最低限の期待値としては妥当な見方でしょう。ただし、このモデルは特に「エージェント型コーディング」において能力を向上させ、概念の明確な特定や科学的研究、事実の正確性においても優れた成果を示しています。
関連記事:GPT-5.5 を 10 ラウンドにわたってテストした結果、93/100 のスコアを獲得。高揚感のみで減点されました
なぜ重要なのか:モデル自体が直前のバージョンから劇的に進化したわけではありませんが、5.4 から 5.5 への更新がわずか 2 か月未満で行われたことは、エージェント型コーディングの進展が OpenAI のモデルリリースサイクルをいかに加速させているかを物語っています。ZDNET のデイヴィッド・ゲワルツ氏が解説しているように、同社は AI を用いて AI を構築する他の最先端研究機関と同様に、更新頻度を指数関数的に高めています。
ChatGPT Images 2
OpenAI | 2026 年 4 月 23 日
何をするものか: Sora(生成動画モデルおよびソーシャルプラットフォーム)の終了直後、OpenAI はやや混乱を招く形で「Images 2」を発表しました。ZDNET のモデルテスターであるデイビッド・ゲヴィッツ氏は、リリース前に Images 2 を早期に体験し、その性能に感銘を受けました。彼は正式な「エキスパートスコア」は付与していませんが、このモデルは「楽しい」「飛躍的な進化だ」と評価し、「実際の業務でも有用だ」と述べています。
なぜ重要なのか: OpenAI は Sora を終了させたことで、消費者向けの AI 製品から撤退する動きを見せていました。その背景には、Lucrative な企業契約の獲得において Anthropic に先を越されたという事情があります。しかし、OpenAI がこの方向転換の文脈の中で依然として Images 2 を発表したことは、画像生成モデルがエンタープライズ AI にとって十分に重要であると見なしていることを示しています。特に直近で登場した「Anthropic の Claude Design(Figma に対抗するツール)」の影響も無視できません。
Claude Opus 4.7
Anthropic | 2026 年 4 月 16 日
何ができるか: Opus 4.6 の発表から比較的短期間で登場したこのモデルは、誠実さの向上と、同調バイアスやハルシネーション(幻覚)の減少において新たな記録を達成しました。また、サイバーセキュリティ分野での能力も顕著で、このモデルのリリース直後に公開された「Claude Security」をサポートしています。ただし、多くの予想とは異なり、これは Mythos というモデルではありません。
関連記事: Anthropic の新ツール「Claude Security」がコードベースから脆弱性をスキャンし、優先すべき修正箇所を提案
なぜ重要か: ハルシネーションと誠実さは、最高峰のモデルであっても解決が難しい課題です。安全性に真剣に取り組む AI ラボである Anthropic がこれらの分野で大幅な進歩を達成したことは、決して簡単なことではありません。
Claude Mythos(プレビュー)
Anthropic | 2026 年 4 月 7 日
何ができるか: これは少し複雑です。なぜなら、Mythos は実際には一般公開されていないからです。Anthropic は、この新しい汎用モデルが通常通りリリースするにはあまりにも強力だと判断し、大きな話題を呼びました。このモデルは以前の Anthropic モデルから飛躍的な進化を遂げていますが、特にセキュリティ上の脅威となる点に懸念を抱き、以下のように述べています。「コンピュータセキュリティのタスクにおいて、驚くほど高い能力を発揮します。」
これを受けて、Anthropic は Google や Nvidia、Microsoft といった競合他社の AI ラボや、Palo Alto Networks などのセキュリティ当局と協力して「Project Glasswing」を主導しました。この取り組みは、「世界で最も重要なソフトウェアの保護を支援し、サイバー攻撃者に対抗するために業界全体が採用すべき準備を整えること」を目的としています。
関連記事:Apple、Google、Microsoft が Anthropic の Project Glasswing に参加し、世界の最重要ソフトウェアを守る
なぜ重要なのか: Anthropic のガイダンスに従えば、Mythos は世界全体のソフトウェアに対する重大な脅威となります。そのため、アクセスできるのは限られたパートナーのみです。現状のサイバーセキュリティ体制では、モデル能力が急速に進化する最前線に対応しきれていない可能性があります。Mythos が唯一の存在ではなく、他のラボも同様のブレークスルーを達成した際、その後の数多くのモデルの一つに過ぎないかもしれません。
現時点ではリリースからわずか数週間ですが、Mythos はすでに大量のソフトウェアバグを検出する役割を果たしています。
GPT-5.4
OpenAI | 2026 年 3 月 5 日
何ができるか: OpenAI は、GPT-5.2 からわずか 3 ヶ月後にリリースされたこの新モデルを、プロフェッショナルな業務に特化して設計されたと位置付けています。同社自身のテスト結果(第三者による検証が行われるまでは鵜呑みにしないよう注意が必要です)によると、GPT-5.4 は人間の専門家と同等かそれ以上の成果を 83% のケースで達成しました。
なぜ重要なのか: AI 企業は、エンタープライズからの信頼獲得(および契約締結)に注力し、エージェント型 AI がもたらす可能性を称賛する一方で、リスクや遅延、あるいは高すぎるコストを抑えつつ複雑な業務タスクを処理できるモデルが求められています。プロフェッショナルなワークフローでの能力を示すモデルの進展は、AI 導入に苦戦する企業にとって真剣に検討される機会が高まりますが、シームレスな統合を保証するものではありません。
関連記事: OpenAI の新 GPT-5.4、テストで専門レベルの業務において人間を 83% の確率で上回る
Claude Opus 4.6
Anthropic | 2026 年 2 月 5 日
何ができるか: このモデルは、特にコーディング分野において自律的なエージェント業務の基準を短期間で再定義しました。プログラミングタスクに特化して構築する能力で権威を持つ Anthropic の成果である以上、これは驚くべきことではありません。Opus 4.6 はまた、全体的に複雑で実行時間が長いタスクにおいても改善を示しました。
なぜ重要なのか:Opus 4.6 は、タスクをより自律的に処理できるため、ワークフローのより多くの部分を信頼して任せることが可能になります。これは通常、エージェント型サービスが苦手とする領域です。
関連記事:Anthropic の新モデル「Claude Opus 4.6」は、最初の試行で成果物を完璧に仕上げる
GPT-5.3-Codex
OpenAI | 2026 年 2 月 5 日
何ができるのか:この新しいコーディングモデルは、OpenAI が「自身を構築・デバッグする際にも役立った」と述べている通り [1]、タスクの最中に中断して方向転換できるのが特徴です。もしこれが事実であれば、試行錯誤を繰り返す複雑なプロジェクトや、要件が変化するプロジェクトに取り組む開発者にとって大きな恩恵となります。また、GPT-5.3-Codex は実行時間が 24 時間以上持続でき、ユーザーの意図をより深く理解できる点も強みです。
関連記事:OpenAI の新モデル「Spark」は GPT-5.3-Codex より 15 倍高速にコードを書くが、代償がある
なぜ重要なのか:OpenAI は、エージェント型コーディングにおいて Anthropic のリードを追いかけようとしています(偶然か否かは別として、GPT-5.3-Codex の公開日は Anthropic が Opus 4.6 を発表した日と一致しています)。ZDNET の専門家は「バイブコーディング」においては Claude Code を他のツールよりも好む傾向がありますが、OpenAI が噂される通り企業顧客へ重点を移し、娯楽的な消費者向けツールから撤退する方針へと転換すれば、最終的にはその差が埋まる可能性があります。
原文を表示

*Follow ZDNET: *Add us as a preferred source* on Google.*
AI labs are shipping new models nonstop. Besides being better and faster than their predecessors, not every new model is guaranteed to be a major step change, despite how the company's PR may wax poetic about them. Model strengths really emerge in context: Where are competitor models lacking or excelling? Which models have outstanding specialties, and which are just catching up to industry standards?
Also: How we test AI at ZDNET
Our Model Release Tracker helps you make sense of where models stand relative to each other and whether they're worth a deeper look. While we don't test every model or model update on this list, we'll always include the key elements you need to know, along with our hands-on expert test, where applicable. We also include an Expert Score for certain models. Curious about how we test AI? Check out this breakdown of our process.
Here are the biggest model releases of 2026 so far and what to know about them. We'll update this list whenever a notable new model arrives.
Nemotron 3.5 Lightning
Nvidia | Aug. 11, 2026
What it does: Nvidia framed its new model, part of its existing Nemotron 3 family, as a go-to workhorse for "high-volume" agentic tasks. Faster than similar models and customizable, it's geared toward specialized jobs, which it can target more directly alongside NeMo Switchyard, a new routing library Nvidia released with the model. With "the knowledge of Ultra packed into the size of Nano," as generative AI VP Kari Briski said during a briefing, Nemotron 3.5 Lightning can also run locally for better security.
Also: As Meta fades in open-source AI, Nvidia senses its chance to lead
Why it matters: Timing is everything with this release. Nvidia and even proprietary-first AI companies like OpenAI are pushing open-weight models more than ever, amidst recent security incidents. Meta also just released Muse Glimmer, an open agentic model for local use on individual Macs and PCs.
Especially with lauded Chinese startup DeepSeek raising prices on open models, Nemotron 3.5 Lightning is landing in an industry environment that is more amenable to open models than it perhaps ever has been. Nvidia is framing the model's customizability as a security and privacy perk rather than the risk it's been framed as thus far.
Muse Code and Muse Spark 1.2
Meta | Aug. 5, 2026
What it does: Muse Code (powered by the new Spark 1.2) is Meta's first coding agent, finally joining the likes of OpenAI Codex and Claude Code. Based on benchmarks Meta shared, it still lags behind Claude Code and Codex, but occupies the mid-tier of performance. At $1.25 per million input tokens, Muse Spark 1.2 is priced at a fraction of Anthropic's Claude Opus 5 (another strong coding model), which is $5 per million input tokens.
Also: Claude's Record-a-Skill cut my research from hours to 30 minutes - but the magic has limits
Why it matters: Mostly, this is a one-to-watch release. As sophisticated models stay expensive to use and open-weight or good-enough models become more popular, Muse Code may not need to beat Codex or Claude Code to compete. As Meta's first foray into this part of the industry, it's worth watching how developers react to the price trade-offs.
MAI-Cyber-1-Flash
Microsoft | July 27, 2026
What it does: Microsoft's latest model exists inside MDASH, its agentic security hub designed to triage issues. The company said Cyber-1-Flash is built to detect "challenging vulnerabilities in complex codebases" and costs half of what leading models charge. The model scored more than 10 percentage points above Anthropic's Mythos, Google's Gemini, and OpenAI's just-released GPT-5.6 on CyberGym, a security reasoning benchmark.
Also: OpenAI's attack agent did exactly what it was told - just more relentlessly than expected
Why it matters: Given Mythos' storied cyber capabilities -- which drove Project Glasswing and a few government interventions -- Cyber-1-Flash's benchmark performance demonstrates how rapidly AI is evolving overall. The model meets a critical moment in the AI security conversation, where experts' long-held concerns about severe, unpredictable threats are coming to fruition. An OpenAI agent escaped a testing environment and breached Hugging Face, just weeks after an unrelated first-of-its-kind and fully agentic ransomware attack.
Also: Open weights vs. closed: An AI civil war's afoot, and the stakes are existential
Kimi K3
Moonshot | July 16, 2026
What it does: At 2.8 trillion parameters, Kimi K3 is the largest open-source model on the market, designed for "long-horizon coding, knowledge work, and reasoning," according to Moonshot's announcement. For context, that's bigger than DeepSeek V4 Pro, from the company that first made US AI labs nervous back in January 2025, at 1.6 trillion.
Most notably, Kimi K3 topped Anthropic's Fable 5 on the Arena benchmark for front-end coding, which measures complex agentic coding tasks. That said, Moonshot itself noted that Kimi K3's overall performance doesn't top Fable 5, though the model does rival it and OpenAI's GPT-5.6 in several individual benchmarks.
Moonshot released the model's weights on July 27 as promised.
Also: Moonshot AI's new Kimi K2.6 swarms your complex tasks with 1,000 collaborating agents
Why it matters: Kimi K3's competitive performance and record-breaking open-source size reignite debate over whether proprietary American frontier models are worth the cost. Fable 5 is $50 per million output tokens, while Kimi K3 is just $15. Still, open-source models lack the safety guardrails that proprietary models have; those risks are amplified by the fact that Moonshot is a Chinese startup (the US government has been consistently suspicious of Chinese tech).
Inkling
Thinking Machines | July 15, 2026
What it does: Inkling is the first model from Thinking Machines, a startup founded by several ex-OpenAI executives including Mira Murati. It's open-weight and generalist, scoring similarly to other open models on benchmarks for agentic coding and tool use, and slightly better on design and audio. It also used far fewer tokens than major open models like Nemotron 3 Ultra, Thinking Machines said, making it an economical choice.
Why it matters: By starting with an open-weight model, Thinking Machines is characterizing itself somewhat differently than other US-based labs, which prioritize more lucrative proprietary models. The company has also invested in research on designing AI that better keeps humans in the loop during use. While not immediately competitive, Inkling's strong early performance demonstrates a possible open-weight renaissance for the US, as most impressive open models are still China-based.
GPT-5.6
OpenAI | July 9, 2026
What it does: Fresh on the heels of its new voice model, OpenAI released the GPT-5.6 family, which the company said it had cleared with the government a few days earlier. The family includes Sol, a new flagship model, as well as Terra, built for everyday use, and Luna, the economical choice of the three. OpenAI said Sol should optimize cost without sacrificing performance and works with a new "ultra" setting that calls on several agents to speed up task completion.
Sol beat Fable 5, the current darling of the AI race, in adaptive and medium reasoning on UC Berkeley's Agents Last Exam benchmark. Terra and Luna also "outperform Fable 5 at around one-sixteenth the cost," OpenAI said in the release. The model family is about on par with Fable 5 in most major performance areas, including coding, work tasks, and general applications, though it is still faster and cheaper, according to OpenAI.
Also: OpenAI's GPT-5.6 and ChatGPT Work aim to beat Anthropic on price, speed, and productivity
Why it matters: That the government was involved in the timing of GPT-5.6's release signifies a shift -- not in explicit regulation, perhaps, but around a sort of informal partnership between labs and the Center for AI Standards and Innovation (CAIS). Set in motion by the back-and-forth around Anthropic's Mythos and Fable 5, the voluntary testing agreement laid out in the Trump administration's June executive order seems to be going swimmingly.
*(Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)*
Muse Spark 1.1
Meta | July 9, 2026
What it does: An upgrade to Meta's first model, Muse Spark, which it released in April, 1.1 scored above Anthropic's Opus 4.8 and OpenAI's GPT-5.5 -- both impressive models -- for job tasks humans delegate to AI and financial analysis. While the rest of its benchmark results were more scattered, those areas are especially relevant to AI's potential impact on white-collar work and the broader job market.
Also: Meta's new AI tool lets others use your Instagram posts for image generation - how to opt out
Why it matters: Unlike GPT-5.6, Muse Spark 1.1 doesn't rival Anthropic's Fable 5, but it does appear to focus on advancing "personal superintelligence," Mark Zuckerberg's somewhat open-ended niche in the agentic assistant world. For example, Meta's release mentions using Spark 1.1 for "agentic dinner party organization," peeling off from Anthropic and OpenAI's focus on professional use cases and aligning itself more with autonomous capability for personal use, similar to Gemini Spark. Keeping pace with new models from other major labs will allow for better simultaneous testing, which should put the race at an interesting inflection point.
GPT-Live-1
OpenAI | July 8, 2026
What it does: OpenAI said in the release that the new model, which is accompanied by a light version (GPT-Live-1 mini), should make talking to ChatGPT in Voice Mode seem "more like having a real conversation." Because GPT-Live can listen and speak simultaneously, according to OpenAI, that means you should be able to seamlessly interrupt the assistant with questions, pause mid-sentence, and ask ChatGPT to respond more slowly if you need.
OpenAI added that with GPT-Live, ChatGPT Voice can delegate more complex asks to models like GPT-5.5 if Live can't handle them. Everyone can access GPT-Live-1 and GPT-Live-1 mini in ChatGPT for Android, iOS, and in the browser.
Also: I tested ChatGPT's Live Voice upgrade, and it almost felt human - how to try it
Why it matters: While technically a narrower upgrade than a brand-new, highly capable general model, GPT-Live-1 should improve the user experience of talking to AI assistants, an increasingly popular way to engage with them on the go. Eventually, a better voice model could also handle longer, more complex tasks, but we'll have to wait and see.
With the recent rollout of Siri AI, GPT-Live-1 gives users another voice assistant improvement (hopefully) to compare to Siri, though ZDNET's early testing of Siri AI against ChatGPT and Gemini already left something to be desired.
Update: Fable 5 Access
Fable 5 costs more to use, but Anthropic included it as a promotion at no extra cost for paid users, including those on Pro, Max, and Team plans, and****select users on Enterprise plans. Anthropic had originally set Fable 5 to shift to token-based usage on July 7 for paid users but extended paid plan access to July 12.
"As before, you can use up to 50% of your weekly usage limit on Claude Fable 5. After that, you can keep using Fable 5 with usage credits, or switch to another model to keep working within your remaining limits," Anthropic explained in an X post.
After Fable 5 and Mythos 5 were pulled on advisement from the government just days after their release, the government re-allowed access to Mythos 5 for certain partners on June 26. On June 30, the Department of Commerce lifted export controls on both Mythos 5 and Fable 5, and Anthropic began restoring global access to Fable 5 on July 1. Mythos 5 is still only available to specific institutions.
Sonnet 5
Anthropic | June 30, 2026
What it does: Pivoting from its recent focus on Opus models, Anthropic released Sonnet 5. According to Anthropic, Sonnet 5 can "make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models." The company said it performs similarly to Opus 4.8, released a month ago, but costs less. Sonnet 5 starts at $2 per million input tokens but will jump to $3 per million in September. It's now the default in Free and Pro plans and available to all other plan types (Max, Team, and Enterprise).
Also: Why AI tokens will send your enterprise cloud bill sky-high again
In keeping with the agentic focus of the AI industry at large, Sonnet 5 scores notably high on computer use benchmarks and agentic coding -- testing found it completed complex tasks that earlier Sonnet models couldn't. It also comes with safeguards implemented automatically, a nod to both recent PR around Mythos and overall improvements in general intelligence.
Why it matters: With the rate of AI model development continuing to accelerate, it's a fair bet that Sonnet 5 is in fact a significant upgrade over Sonnet 4.6, which was released back in February (eons ago in current shipping time). The timing of this release is also significant; Anthropic noted that Sonnet 5 "has a much lower ability to perform dangerous cybersecurity tasks than our current Opus models," perhaps a stabilizing mention after the complex rollout of its powerful Fable 5 and Mythos 5 models. Ironically, however, Sonnet 5 showed a higher rate of misaligned behavior than Mythos Preview.
Fable 5 and Mythos 5
Anthropic | June 9, 2026
What it does: ZDNET senior contributing editor David Gewirtz called Fable 5 a "defanged" version of Mythos, made safe for public use. Anthropic released Mythos 5 to those who already had Mythos Preview access through Project Glasswing. Fable 5 is restricted from responding to high-risk queries about topics like cybersecurity and biological weapons, but still "Mythos-class," which, by Anthropic's telling, means a step change in capability. As for Mythos 5, Anthropic said it had plans to expand access beyond initial partners through a systematic program. But both models were pulled barely four days after their release on orders from the US government.
Also: Why Anthropic suddenly pulled Fable 5 and Mythos 5 for everyone
Why it matters: The models have created plenty of commotion, not just because of their supposed basis in Mythos and temporary restriction. Fable 5 beguiled safety testers who were unaware that the model was set to downgrade to Opus in response to certain questions, which created trust issues between researchers and Anthropic. Despite being guardrailed, we now know Amazon researchers jailbroke Fable 5 and communicated with the White House to alert Anthropic, though Anthropic framed this incident as "narrow" in its understanding.
While the details remain unclear, it's a rare public instance of government authorities intervening in an American AI product, which could mean Mythos-level models have caused a shift in the Trump administration's relatively hands-off approach to AI labs thus far.
MAI-Thinking-1
Microsoft AI | June 2, 2026
What it does: At its Build developer conference, Microsoft said that this new 35-billion-parameter model is, unsurprisingly, designed for multi-step agentic tasks. It scored similarly on the SWE Bench Pro benchmark test for coding as Anthropic Opus 4.6. The company also noted that enterprise users can trust this model for any use because it was trained only on clean, commercially safe data -- an important tidbit given mounting AI copyright lawsuits.
Also: Microsoft's first reasoning model is one of 7 AIs just released at Build - what we know so far
Why it matters: This is the first reasoning model from Microsoft AI, a notable milestone for any AI lab, but especially so this late in the race. Most labs have created multiple advanced reasoning models at this point, so it's unclear where Microsoft's approach stands relative to those, though the company is focusing on appealing to enterprise clients specifically.
Claude Opus 4.8
Anthropic | May 28, 2026
What it does: Replacing Opus 4.7 starting May 28 (at the same price), Opus 4.8 offers faster thinking modes for one-third the cost of the earlier version, according to Anthropic. Like most of Anthropic's models, 4.8 prioritizes coding abilities, scoring higher than 4.7 on two coding benchmarks but not fully besting OpenAI's GPT-5.5. It also "reaches new highs on our measures of prosocial traits like supporting user autonomy and acting in the user's best interest," the company noted in the release, though definitions for what that means remain murky.
Also: I compared Claude Opus 4.8 with 4.7 in a 10-round honesty test - and a legal prompt broke it
Why it matters: Anthropic has always prioritized model safety and interpretability, but appears to be further emphasizing that standard with this release. The company said Opus 4.7 had a 92% honesty rate, in addition to being less sycophantic and hallucination-prone overall. The fact that it claims 4.8 shows "substantially" lower rates of misalignment than 4.7 indicates an increasingly high standard for model safety, especially because Anthropic compared 4.8's alignment to that of Mythos Preview.
Gemini 3.5 Flash
Google | May 19, 2026
What it does: At Google I/O, the company launched its 3.5 model family, designed for building agents (following suit with where all of the AI industry is focused at the moment). The first model available now is Gemini 3.5 Flash, which beat Gemini 3.1 Pro on a few coding and agentic benchmarks and now undergirds AI Mode in Search and the Gemini App as the default model. Like all Flash models, it's optimized for speed and a cheaper, more lightweight user experience, but can still handle "long-horizon" agentic tasks, Google noted. 3.5 Pro should roll out in June.
Also: Anthropic launches Opus 4.8, with honesty as its killer feature
Why it matters: The most important place 3.5 will operate is in Search. As Google doubles down on an AI-first search experience, the models running it shape the responses people will passively consume every day, meaning it has to be especially up to snuff in terms of accuracy and hallucinations. Notably, though, the System Card for 3.5 Flash lacks any mention of hallucination rate or sycophancy.
GPT-5.5 Instant
OpenAI | May 5, 2026
What it does: OpenAI said in its announcement that the lighter version of OpenAI's just-released GPT-5.5 is less verbose than its predecessor, GPT-5.3 Instant. It also touted fewer hallucinations and improved factuality, saying "GPT‑5.5 Instant produced 52.5% fewer hallucinated claims than GPT‑5.3 Instant on high-stakes prompts covering areas like medicine, law, and finance."
Also: Anthropic's Mythos is evolving faster than expected, reports AI safety agency
Why it matters: GPT-5.5 Instant replaces GPT-5.3 as the default model in ChatGPT. Again, while the expectation is that each new AI model gets more efficient, gets easier to use, and makes up less stuff, a significant improvement in hallucinations for a model most people use for fast queries could mean less misinformation spreading among the masses. That's especially critical given how many people are using ChatGPT for everyday health questions, for example.
*(Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)*
Nemotron 3 Nano Omni
Nvidia | April 28, 2026
What it does: The latest in Nvidia's open Nemotron family, this model provides agents with multimodal input. That means they can "perceive and reason across visual, audio, and textual inputs within a single shared perception‑to‑action loop," according to Nvidia, thereby unifying multiple capabilities into a single system.
Also: AI is an arms race, and the US wants $9 billion in Nvidia superchips to keep up
Why it matters: Normally, systems of agents need to use separate models for speech, vision, and text, meaning they jump across documents, video, and audio to complete multi-step tasks. That slows down workflows, undermines the context agents gather, and racks up inference costs. Nvidia's approach, if it works, would streamline this process and reduce token use, saving you money. Try it on Hugging Face.
GPT-5.5
OpenAI | April 23, 2026
Expert Score: 93/100
What it does: ZDNET tester-in-residence David Gewirtz technically gave GPT-5.5 an A- score, but said it "can be reductively described as better and faster than GPT-5.4," which is hopefully the bare-minimum expectation for a new model. Specifically, though, the model got better at agentic coding, clearly identifying concepts, scientific research, and factual accuracy.
Also: I put GPT-5.5 through a 10-round test: It scored 93/100, losing points only for exuberance
Why it matters: While the model itself may not be leaps and bounds ahead of its immediate predecessor, the quick turnaround from 5.4 to 5.5 -- less than two months -- indicates how rapidly agentic coding is accelerating OpenAI's model release cycle. As ZDNET's David Gewirtz breaks down, the company, much like other frontier labs using AI to build AI, is shipping updates at an exponentially increasing rate.
ChatGPT Images 2
OpenAI | April 23, 2026
What it does: Soon after sunsetting Sora, its generative video model and social platform, OpenAI somewhat confusingly announced Images 2. ZDNET model tester David Gewirtz got an early look at Images 2 before its release and was impressed. While he didn't give this model a formal Expert Score, he said it's fun, a huge leap, and actually useful for work.
Why it matters: OpenAI seemed to be getting out of the more consumer-minded AI product game when it discontinued Sora, having been beaten by Anthropic at securing lucrative enterprise contracts. That OpenAI still came out with Images 2 within that redirection narrative indicates that it sees image generators as relevant enough to enterprise AI -- especially on the heels of Anthropic's Claude Design.
Claude Opus 4.7
Anthropic | April 16, 2026
What it does: Arriving relatively quickly after Opus 4.6, this model boasts new highs in honesty and reductions in sycophancy and hallucinations. It also appears to have a knack for cybersecurity, as it backs the new Claude Security, released shortly after the model itself -- but no, it's not Mythos, as many suspected.
Why it matters: Hallucinations and honesty are among the most difficult issues plaguing even the best models. For Anthropic to claim such significant gains in those areas is no small feat for an AI lab that takes safety seriously.
Claude Mythos (Preview)
Anthropic | April 7, 2026
What it does: This is a tough one because Mythos isn't actually available to the public. Anthropic created quite a media storm when it positioned the new general-purpose model as too powerful to release as usual. While the model is apparently a step change from earlier Anthropic models, the company was especially alarmed because of the security threat it posed, stating that "it is strikingly capable at computer security tasks."
In response to that, Anthropic spearheaded Project Glasswing, a collaborative effort with several rival AI labs, including Google, Nvidia, and Microsoft, as well as security authorities like Palo Alto Networks, "to help secure the world's most critical software, and to prepare the industry for the practices we all will need to adopt to keep ahead of cyberattackers."
Why it matters: If we're to believe Anthropic's guidance that Mythos poses a significant threat to the world's software -- so much so that only a select few partners can access it -- cybersecurity apparatuses as they stand may not be prepared to meet the rapidly evolving frontier of model capabilities. Mythos may not be the only model of its caliber but may simply be the first of many to come once other labs achieve similar breakthroughs.
For now, just a few weeks into its release, Mythos is helping catch software bugs in droves.
GPT-5.4
OpenAI | March 5, 2026
What it does: OpenAI framed this new model, released barely three months after GPT-5.2, as specifically designed for professional work. According to the company's own testing (which should always be taken with a grain of salt until verified by a third party), GPT-5.4 matches or outperforms human professionals 83% of the time.
Why it matters: As AI companies focus more on gaining enterprise trust (and contracts) while lauding what agentic AI can do, they need models that can handle complex work-related tasks with minimal risk, delay, or prohibitively high costs. Any model advancement that shows prowess in professional workflows has a better chance of being taken seriously by companies struggling to adopt AI, though nothing guarantees seamless integration.
Also: OpenAI's new GPT-5.4 clobbers humans on pro-level work in tests - by 83%
Claude Opus 4.6
Anthropic | Feb. 5, 2026
What it does: This model quickly redefined the standard for autonomous agentic work, especially for coding. That's no surprise given Anthropic's authority in building models especially adept at programming tasks. Opus 4.6 also demonstrated improvement in complex, longer-running tasks overall.
Why it matters: Opus 4.6's ability to handle tasks better on its own means you can reliably offload more of your workflow to it -- something agentic offerings usually struggle with.
Also: Anthropic says its new Claude Opus 4.6 can nail your work deliverables on the first try
GPT-5.3-Codex
OpenAI | Feb. 5, 2026
What it does: This new coding model -- which OpenAI said helped build and debug itself -- can be interrupted and redirected mid-task, which, if true, is a huge boon for developers using it on complex or shifting projects with tons of trial-and-error. GPT-5.3-Codex also boasts run times of over a day and a better grasp on user intent.
Also: OpenAI's new Spark model codes 15x faster than GPT-5.3-Codex - but there's a catch
Why it matters: OpenAI is trying to catch up to Anthropic's lead in agentic coding (and, coincidentally or not, released GPT-5.3-Codex on the same day as Anthropic launched Opus 4.6). While ZDNET experts often prefer Claude Code to other tools for vibe coding, OpenAI's rumored shift toward enterprise clients and away from fun consumer tools could eventually close that gap.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み