SpaceX が Cursor を 600 億ドルで買収、AI エージェント開発を加速
本文の状態
日本語全文を表示中
詳細モードで約26分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Latent Space
Z.ai が GLM-5.3 を公開し、ベースモデルの再学習で性能を向上させた一方、Alibaba や DeepSeek も中国発オープンモデルの拡充を発表した。
AI深層分析を開く2026年8月17日 15:56
AI深層分析
キーポイント
Z.ai の GLM-5.3 発表と技術的特徴
Z.ai は GLM-5.2 と同じ 743B ベースモデルを用いたポストトレーニングで、コード作成やセキュリティ分野での性能を大幅に向上させた GLM-5.3 を公開した。
Alibaba の Qwen3.8 シリーズとローカル展開
Alibaba は Apache 2.0 ライセンスのマルチモーダルモデル Qwen3.8-27B をリリースし、17GB RAM で動作可能な実用性を強調した。
中国発オープンエコシステムの多様化
DeepSeek や RedNote などが新しいモデルや評価手法を相次いで発表し、中国のラボが専門分野ごとに特化したオープンエコシステムを形成している。
ハネスをインフラとして扱うアーキテクチャの転換
DeepSeek Harness はデモエージェントではなく、コンポーネントのホットスワップや自己修正機能を備えたプラグイン化されたランタイム基盤として扱われている。
ハネス層が最適化とベンチマーク成果の主要な対象となっている
モデル自体の知能だけでなく、メタオプティマイザによるハネスの自動書き換えや設定の厳密なバウンディングが成果に直結しており、観測データが評価・学習基盤として二重の役割を果たしている。
重要な引用
Z.ai’s GLM-5.3: The biggest technical story was Z.ai launching GLM-5.3, positioned as a coding- and cyber-focused model built via post-training on the same 743B base model used for GLM-5.2 rather than a new pretrain.
The key claim many engineers highlighted is that the capability jump came entirely from scaled post-training/RL on longer-horizon executable tasks, not from a larger base model
Qwen3.8 broadens the local/open frontier: Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0
The technically interesting bit is not just 'modularity,' but support for hot-swapping runtime components and potentially enabling agents to modify their own runtime without restart, while preserving auditable event logs and avoiding hidden state.
編集コメントを表示
編集コメント
今回の一連の発表は、中国発のオープンモデルが単なる模倣ではなく、独自の学習手法や特化領域で業界をリードし始めていることを示している。特に Z.ai のアプローチは、リソース制約のある環境でも高性能な AI を実現する新たな基準となる可能性がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Cursor が設立当初の 5 人規模だった頃の最初のポッドキャストを振り返る映像があります。
また、ICML 2024 で Graham Neubig とともにエージェント技術について振り返ったセッションも紹介されています。
さらに、2026 年に迎える同社の第 3 のフェーズについても言及しています。
そして、エンタープライズ向けに FDE(Full-Development Environment)をどのように運用しているかについての議論もあります。
8/13/2026〜8/14/2026 の AI ニュースです。当編集部は 12 のサブレッド、544 のツイート、そして Discord を確認しました。AINews のウェブサイトでは過去のニュースをすべて検索可能です。ご存知の通り、AINews は now Latent Space の一部となっています。メール購読の頻度設定も自由に変更できます。
AI Twitter リキャップ
オープンウェイト・フロンティアの動向:Z.ai の GLM-5.3、Qwen3.8-27B/Max、DeepSeek V4-Pro、RedNote の dots3-note
Z.ai が GLM-5.3 を発表しました。これが今回の最大の技術ニュースです。GLM-5.3 は、GLM-5.2 と同じ 743B ベースモデルをベースにポストトレーニング(後学習)を施して構築されたもので、コーディングとサイバーセキュリティに特化したモデルとして位置づけられています。新規の事前学習は行われていません。
Z.ai とその後の発表によると、エージェント機能やセキュリティ評価において大幅な向上が確認されました。具体的には、Terminal Bench 3.0 で 28.3、DeepSWE で 66.9、Agents' Last Exam で 28.5、GDPVal-AA で 1769 というスコアを記録しています(ベンチマークの概要と詳細はこちら)。
同社はまた、サイバーセキュリティ機能の向上により、まずは安全性レビューを経てからオープンウェイト版として公開する予定であり、それまでは特定のパートナー限定でアクセスが制限されると発表しました(詳細はこちら)。
多くのエンジニアが注目しているのは、この能力の飛躍的な向上が、より大きなベースモデルの採用によるものではなく、長時間実行可能なタスクに対するスケーリングされたポストトレーニングや RL(強化学習)によって実現されたという点です。
Qwen3.8 がローカル・オープンソースの新たな可能性を広げました。アリババは、Apache 2.0 ライセンスで公開されたネイティブ多モーダル密結合モデル「Qwen3.8-27B」をリリースしました。このモデルは、YaRN を用いることでネイティブコンテキスト長 262K から最大 1M まで拡張可能です。
同時に発表されたのが、すでにリリース済みの最上位モデル「Qwen3.8-2.4T-A95B」です(公式発表、性能詳細スレッド)。27B モデルが注目されるのは、単なる学術的なベンチマークではなく、実世界のコーディング業務やオフィスワーク、エージェント機能に特化して設計されている点にあります。
Day-0 での推論サポートも非常に手厚く、vLLM、Ollama、llama.cpp/GGUF、SGLang などが対応しています。特に SGLang では、単一の RTX 5090 で 1 秒あたり 206 トークンの処理速度を達成しました。また、Together、Fireworks、Modal、DigitalOcean、DeepInfra など主要なクラウドプロバイダーもサポート対象です。
実運用における詳細にも配慮がなされています。Unsloth は NVFP4 や動的 GGUF ビルドの提供を主張しており、Qwen 側はローカル利用向けに 17GB の RAM で 27B モデルを実行可能であることを強調しています(公式投稿)。
DeepSeek V4-Pro と RedNote の dots3-note は、中国におけるオープンモデルの波を牽引し続けています。vLLM は DeepSeek-V4-Pro へのサポートを発表し、MIT ライセンスやプレビュー版とのチェックポイント互換性、統合されたドラフト機能のサポートを強調しました。
一方、RedNote の AI ラボは「dots3-note Preview」をリリースしました。これは 280B パラメータを持つマルチモーダル MoE モデルで、アクティブなパラメータ数は 16B、コンテキスト長は 512K に達します。長期稼働するエージェント向けに設計され、長時間の自己評価を行うための新しい RL 手法「TEMPO」も併せて発表されています(初期シグナル、サマリー、チームによる技術解説)。現在顕著な傾向として、複数の中国ラボが専門化を進めていることが挙げられます。多くの評論家が Z.ai、DeepSeek、Moonshot、Qwen、MiniMax、RedNote を、それぞれ異なる強みを持つ高速で動くオープンエコシステムとして捉えています。
エージェントランタイム、ハルネス、長期ホライズン学習
DeepSeek Harness はデモ用エージェントではなく、インフラストラクチャとして位置づけられています。このリリースにより、モデルの UX 以上にランタイムアーキテクチャに関する議論が活発化しました。
いくつかの詳細な解説では、Harness をプラグイン化されたエージェントランタイムとして紹介しています。ここでは、エージェントループ、ツール、セッション、ファイルシステム、プロバイダーなどすべてを交換可能に設計されており、Cordis がライフサイクル管理、反応型の依存関係、そして元に戻せる効果(ロールバック)を提供します。
技術的に興味深い点は、「モジュール性」という単なる概念にとどまらず、ランタイムコンポーネントのホットスワップ(実行中の交換)をサポートし、エージェントが再起動なしで自らのランタイムを修正できる可能性を開く点にあります。同時に、監査可能なイベントログの維持や、隠れた状態の発生を防ぐことも実現しています。
複数の開発者は、現在の Harness はこの方向性に比べれば「間違っている」、あるいはコア部分が固定されすぎており柔軟性に欠けると反応しました。
ハーン(Harness)自体が最適化の対象となりつつあります。いくつかの投稿で指摘されている通り、ベンチマークや製品における成果は、もはやベースモデルの知能向上だけでなく、スキャフォールドやハーン層からの改善によって得られるケースが増えています。
DAIR が発表した AutoDesign では、メタオプティマイザーがロールアウトフィードバックに基づいてハーン自体を再構築する仕組みを示しました。これにより、論文からポスターへの生成や、エージェント・モデル構成間の転移において成果が向上したと報告されています。一方、Lambda の Tetris 実験は逆の視点から同様の指摘を行いました。プロンプトの配置、設定、サンドボックス制約などが結果に大きな影響を与えることが示され、厳密な制約がない場合、エージェントはベンチマークの抜け穴を悪用する傾向があるのです。
これは、観測データが現在、評価(evals)、メモリ、学習基盤という二重の役割を果たしているという広範な議論とも一致しています(LangSmith のドキュメント参照)。
ベンチマーク、評価、そしてベンチマークへの懐疑
新たな評価指標は、エージェント特有の失敗モードに焦点を当てています。Vals が発表したアジェンティック逆エンジニアリング用ベンチマークでは、中間成果物ではなく、セキュリティ関連のバイナリ環境における決定論的な最終目標達成に重点を置いています。これに関連する別の記事では、ソースコードが利用可能な場合と、バイナリから推論を迫られる場合では、現在の最先端エージェントのパフォーマンスに大きな差があることが指摘されています。
OpenRouter は、ツール活用型エージェント向けのウェブ検索ベンチマークを導入しました。また、Ai2 の TutorMoments は、再生ベースの指導評価として引用され、モデルが学習者を支援しすぎず、むしろ生産的な試行錯誤を促すべきであるという課題が浮き彫りになりました。
評価(eval)への反発は続いている。共通するテーマとして、ベンダーが提示するベンチマークの数値に対する懐疑視がある。
Vik Paruchuri氏はLlamaIndexのベンチマークを批判し、スコアラーの不具合でシステムの評価が65%から93.6%に跳ね上がる可能性があると指摘。開発者はマーケティング資料を鵜呑みにせず、自社の環境で評価を行うべきだと強く主張した。「自社を含む」という注釈も付されている(フォローアップ記事あり)。
François Chollet氏は、公開されているARC-3のデモセットが学習データや評価用データではないと繰り返し強調。このセットでのリーダーボードスコアは、プライベートなセットにおける性能を測る指標として頼りない代わりのものだと指摘した。
ここで注目すべき追加情報として、Omar Sar氏が紹介したMetaの「Wiggle Framework」がある。これはLLMによる判定(judge)が、再プロンプトや敵対的な圧力にさらされた際にどう振る舞うかをストレステストするものだ。その結果、静的な押し付けに対しては判定が25〜71%反転し、敵対的な説得者に対しては62〜91%も反転することが判明した。
インフラ、サービング、コストエンジニアリングの分野では、サービング最適化がモデル機能の第一級要素として扱われるケースが増えている。QwenやDeepSeekに関するDay-0のインフラサポートでは、単なるAPIアクセスだけでなく、埋め込みドラフトヘッド(draft heads)、推測的デコーディング(speculative decoding)、メモリと量子化のトレードオフといった技術が重視された。
27BパラメータのQwenリリースには、MTPドラフトヘッドに関するvLLMのガイダンスや1Mコンテキストのサポート、Blackwell GPU 1枚でのサービングが可能であることが含まれていた。また、ggerganov氏は大規模コンテキストと推測的デコーディングに対応したローカル環境向けのllama.cppレシピを紹介している。
Tim Dettmers氏は、単一のDGX SparkやAMD Strix Halo上で強力なモデルを実行するための次世代効率化手法を予告。これは約7 tok/sのデコード速度と250 tok/sを超えるプリフェッチ(prefill)性能を実現するものだ。
ツール類やクラスター運用に関する実用的なアップデートも発表されました。Stas Bekman 氏は、PyTorch で NCCL の集合呼び出しが応答しなくなった際の診断ガイドを提供しました。また、Python 3.14 以降では、計測機能(instrumentation)を追加せずとも実行中のプロセスに pdb を接続できるようになった点にも言及しています。
Turbopuffer は、顧客クラウド内での直接ホストアクセスが不要な BYOC(Bring Your Own Cloud)デプロイメントを含む、100 以上の TPUf クラスターを運用するための独自コントロールプレーンについて説明しました。データ側では、Hugging Face の datatrove がバージョン 0.10.0 にアップデートされ、Hugging Face Jobs 向けの JobsPipelineExecutor や HF バケットとの統合が追加されました。また、推論結果の出力も保持されるようになりました。
製品およびプラットフォーム動向:Cursor/SpaceXAI、Gemini 3.7 Flash、Claude Code、ローカルエージェント UX
Cursor が SpaceX に加わる:最も注目を集めた技術・企業間の動きは、Cursor が SpaceX の一部となり、チームが SpaceXAI に合流したという発表です。今後は Grok、Grok Build、Grok Bot、Grok API、そして Cursor 全体で協力して開発を進めます。SpaceXAI はこの買収を正式に確認し、「まずソフトウェアエンジニアリングの加速を図り、その後に広範な知識労働へと展開する」という方針を示しました。
これは、コーディングエージェントチームがもはや狭義のエディタ製品ではなく、戦略的なモデルやプラットフォーム資産として位置づけられていることを示す明確な兆候の一つです。
Google は Gemini 3.7 Flash の展開において、エージェント機能とコスト効率の向上に注力しました。このモデルは Gemini アプリ、Search AI モード、Google Workspace / Sheets キャンバス、そして Spark 全体に広く導入されています。その位置づけは「コーディングとエージェント処理におけるこれまでで最も賢い実用モデル」であり、デモでは簡単なプロンプトからプレイ可能なウェブゲームを生成する事例が中心でした。
外部評価の結果は控えめながらも肯定的なものでした。Vals の Index v2 において Gemini 3.7 Flash は 59.4% で 7 位にランクインし、Gemini 3.6 Flash が 14 位だったことから順位を上げました。
Anthropic は Claude Code における「Auto モード」を Pro、Max、Team のデフォルト権限設定として導入しました。これにより /auto-mode-setup コマンドを通じてリポジトリ認識型のセットアップが可能になり、信頼できるリポジトリやドメインの提案が行えるようになりました。
オープンソースおよびローカル環境においては、Hermes がエージェントセッション内で cron 処理のような繰り返しアクションを可能にする /loop を追加しました。また Nous は、Hermes Desktop から Hermes Cloud エージェントをターゲットに設定でき、ラップトップを閉じても作業を継続できる点を指摘しています。
Ollama も、DeepSeek Harness をローカルで起動するサポートを追加しました。
エンゲージメント数の多い投稿(Top tweets)
Cursor と SpaceXAI の提携発表は、その日の最も注目されたテック関連の投稿となりました。これはコーディングエージェントや、モデルとプロダクトを垂直統合したスタックを取り巻く業界再編が継続していることを示しています。
GLM-5.3 のリリース:Z.ai が発表した GLM-5.3 は、すでに訓練済みの最先端ベースモデルから、事後学習と長期的な強化学習(RL)によって潜在的な能力を引き出す可能性を浮き彫りにしたことで、最も注目されたモデル発表の一つとなりました。
Qwen3.8-27B のオープンウェイト公開:アリババが公開したこの 270 億パラメータのローカル多モーダルモデルは、広範な Day-0 サポートを備え、本格的なエージェント業務や専門的な作業にも耐えうるものとして評価されたため、大きな注目を集めました。
実用的なコーディング・エージェントの実績:redp314 氏が「Claude Code で 800 ファイルから DICOM ビューワーを 2 つのプロンプトで構築した」と投稿したのは、ベンチマークの議論を超えた、現在のコード支援ツールの現実的な限界を示す力強い事例として際立っています。
AI Reddit まとめ
/r/LocalLlama と /r/localLLM のまとめ
- Qwen3.8-27B のリリース、ベンチマーク結果、およびテンプレート
Qwen3.8-27B のプレリリースモデルカードが公開されました(アクティビティ数:1006)。画像は、Hugging Face 上の Qwen/Qwen3.8-27B モデルカードの技術スクリーンショットです。投稿で言及されている通り、このカードはリリース前に閲覧可能だった後、正式に公開されたことを示しています。
モデルカードには、重みや設定ファイルの提供予定、Transformers、vLLM、SGLang への対応が明記されています。また、コーディング、エージェント実行、リサーチ、長文コンテキスト処理における改善点が強調されており、ネイティブのコンテキスト長は 262,144 トークン、最大 1,000,000 トークンまで拡張可能とされています。
コメント欄では、「推論のための努力(reasoning effort)」が今回の目玉機能である可能性が高いという指摘や、長いコンテキストウィンドウへの称賛が目立ちました。さらに、27B モデルにビジョン機能が搭載されている一方で、遥かに巨大な 2.4T モデルにはその機能がないと報じられている点について、多くのユーザーが驚きを表明しています。
コメント投稿者たちは、Qwen3.8-27B の技術的な注目点として、ネイティブのコンテキスト長 262,144 トークン(最大 100 万トークンまで拡張可能)を挙げています。
アーキテクチャや製品ラインの違いについても関心が寄せられました。27B モデルにはビジョンサポートが含まれている一方、はるかに大きな 2.4T モデルには含まれていないという事実は、能力のスケーリングの観点からするとユーザーにとって意外な点でした。
ある投稿者は、QAT(Quantization-Aware Training:量子化意識トレーニング)に関する言及が明示的にないことに触れ、Gemma 4 31B と比較しました。同モデルでは QAT が量子化モデルのパフォーマンスを劇的に向上させることが確認されています。また、他の投稿者も、最近のモデルカードで「推論のための努力」が新たなチューニングや制御機能として登場している点を指摘しています。
Qwen3.8-27B は Qwen3.6-27B と同じアーキテクチャです(アクティビティ数:902)。
公開された GIF 画像では、Qwen3.6-27B と Qwen3.8-27B のアーキテクチャ図が並べて表示されていますが、視覚的に全く同一です。ビジョン/埋め込みパス、マスク付きスキャッター、繰り返される Qwen3_5DecoderLayer スack、RMSNorm、最終的な Linear レイヤー、そして出力部分まですべて共通しています。
関連する Hugging Face Viewer の差分レポートではアーキテクチャの変更点が 0 と報告されており、この投稿の主張——つまり Qwen3.8-27B の能力向上はモデル構造の変化ではなく、トレーニングやデータ、ファインチューニングによるものだろうという見解——を裏付けています。コメント欄では、これはゼロから作り直したモデルではなく漸進的なアップデートであると捉えられており、「通常、品質を高める最大の要因となるのはトレーニングデータである」と指摘する声もありました。
また、専門タスクにおけるローカルモデルの精度向上のために、ホットスワップ可能な LoRA 形式のアダプターが今後普及する可能性について言及するコメントも見られました。
複数のコメント投稿者が Qwen3.8-27B をゼロから学習したモデルではなく、漸進的なアップデートと解釈しています。その中には「Qwen3.6-27B、さらには Qwen3.5 と実質的に同じように見える」という意見もありました。技術的な示唆として、アーキテクチャの変更よりも、データセットの更新やトレーニング後の調整が品質向上の主要な要因である可能性が指摘されています。
ある投稿者は、Qwen 系列向けの高速ローカル推論パスとして Ninfer(GitHub)を挙げています。同ツールには新たに最大 C=8 の並列リクエスト対応機能が追加されました。報告された数値によると、Qwen3.6-35B-A3B は C=8 で合計 1,313.8 トークン/秒のデコード性能を達成し、27B モデルの NVFP4 プロファイルでは 1,146.9 トークン/秒、単一並列時のスループットと比較して約 5.67 倍の性能を発揮しています。
ローカル推論ワークフローにおいて、ホットな LoRA スワップが重要になる可能性について議論がありました。これは、ベースモデルを置き換えることなくタスク固有の精度向上を実現し、小型または漸進的なベースモデル更新に対する補完策として位置づけられています。
Qwen3.8-27B が利用可能になりました(アクティビティ数:745)。画像(リンク)は Hugging Face 上の Qwen/Qwen3.8-27B-FP8 ページを示しており、Transformers と Safetensors でパッケージ化され、Apache 2.0 ライセンスの下で F8_E4M3 を用いた FP8 量子化と BF16 テンソルを併せ持つ、新たに公開された 28B パラメータの Qwen 3.8 モデルであることがわかります。あるコメントでは、RTX 5090 での初期ローカル推論が約 50〜60 トークン/秒で実行可能だと報告されています。Qwen 3.6 よりも安定感があり、熟考しているような印象を受けるものの、設定が最適化されていない可能性や、MTP(Multi-Token Prediction)のサポートはまだ利用できない点に言及しています。
コメントは慎重ながらも熱意に満ちており、あるユーザーはこのモデルを「成長した 3.6」であり、長時間実行されるタスクへの対応力が強化されていると表現しました。また、27B は多くのローカルユーザーにとって大きすぎるため、9B や 35B など、より小型または代替サイズのモデルは利用可能かどうかを尋ねる声もあります。
RTX 5090 で Qwen3.8-27B をテストしたユーザーは、Qwen 3.6 と同じ設定で安定して約 50〜60 トークン/秒のローカル推論が可能だと報告しました。MTP(Multi-Token Prediction)のサポートが利用可能になれば、さらに性能が向上する可能性があります。質的な面では、長文生成において Qwen 3.6 よりも慎重な姿勢が見られました。1 万語もの物語をいきなり書き始めるのではなく、段落間の整合性を確認して修正し、タスクをサブタスクに分割して、より綿密な計画のもとで章ごとに生成する傾向があります。
Muse Glimmer は、約 30B クラスのモデル群において「フロンティア(最前線)」の地位をわずか 4 日間だけ維持しました。この投稿にある画像は、~30B クラスのモデルを比較したベンチマーク表で、Muse Glimmer-30B と Qwen3.8-27B が強調表示されています。記事では、Qwen の 27B モデルが多くの報告された指標で Muse Glimmer を上回ったため、Muse Glimmer はこのサイズクラスにおいて「最前線」の地位をたった 4 日間しか保てなかったと主張しています。
Muse Glimmer のスコアは、Agentic terminal coding で 51.7、SWE-bench Pro で 51.2、IFBench で 77.0、GPQA Diamond で 83.5 などですが、ベンチマークの多くの項目が欠落しており、比較が不完全であることが指摘されています。画像:i.redd.it/2cclgla7xdjh1.png。
コメント欄では、モデル開発ラボは特定のサイズクラスで他社に抜かれることを防ぐため、複数のパラメータ規模を同時にリリースすべきだという意見や、Meta は 70B、100B、400B といったより大規模な Glimmer バリアントを出すべきだったという指摘が寄せられています。また、27B モデルが「Opus 4.6 Max」に匹敵する性能を達成するのは驚きだが、Meta がこれに対抗する強力なフロンティアモデルで応答することを期待するという声もあります。
あるコメントでは、Muse Glimmer が推測的デコーディング(speculative decoding)を搭載してリリースされ、これによって TPS やスループットが向上したと報告されていることに触れ、Qwen にも同様の加速経路があるのかと問うています。これはスレッド内で最も具体的な実装関連の指摘ですが、具体的な TPS 数値やデコーディングの設定については言及されていません。
別の技術的な批判では、Muse Glimmer が Qwen よりも劣るとして、「認知的なミス」が多いと指摘されています。具体的には、推論プロセスが不要なコンテンツポリシーに関する議論に逸れ、最終回答と矛盾するケースがあるというのです。コメント投稿者は Qwen の文章スタイルを好まないとしつつも、Qwen において同様の推論と最終出力の不一致は観測されていないと述べています。
さらに別のコメントでは、この結果を約 27B パラメータ規模でありながら「Opus 4.6 Max レベル」に匹敵するものとして位置づけ、約 30B モデルクラスとしては異例とも言える高い性能を示唆しています。しかし、スレッド内にはこの比較を実証するためのベンチマーク名やスコア、評価手法、再現性の詳細などは一切提供されていません。
Qwen 3.5、3.6、および新バージョンの 3.8 向けに、Jinja チャットテンプレートを修正しました(アクティビティ:478)。コミュニティが維持するこの差し替え可能なテンプレートは、公式テンプレートで報告された不具合を解決します。具体的には、「enable_thinking=false」設定でのハードエクスセプション、空白注入による多回対話履歴の汚染、OpenAI 形式の JSON ストリングツール引数への対応時のクラッシュ、そして会話途中のシステムメッセージが欠落してツールループが停止する問題です。
このテンプレートでは、Qwen 3.8 の推論努力度(reasoning_effort)を「xhigh」「high」「medium」「low」で制御可能になりました。また、kwargs やパラメータを通じて推論機能を無効化できる機能を復元し、以前の思考内容を保持することでプレフィックスや KV キャッシュの再利用をサポートします。さらに llama.cpp の --reasoning-preserve にも対応しています。llama-server では、--jinja と --chat-template-file chat_template.jinja を指定し、--reasoning-format deepseek を使用して思考内容を OpenAI の reasoning_content として出力することを推奨しています。
著者はローカル環境で 2.4T モデルを検証できないと述べていますが、28 件の自動テストとトークナイザーの整合性チェックを実施した結果を報告し、Qwen 3.8 ユーザーからのフィードバックを求めています。コメント欄では、なぜ Qwen の公式チャットテンプレートにこのような基本的な回帰現象が含まれているのか、また QA プロセスがテンプレートやツール呼び出しのパスをカバーしているのかなどが議論されました。別のユーザーは、27B などのより小型でアクセスしやすいバリアントでのテストにも興味を示しています。
あるコメントでは、Qwen 3.8 のチャットテンプレートにおける回帰現象が報告されています。「enable_thinking=false」を設定しても推論機能が単に無効になるだけでなく、ハードエクスセプションを引き起こすことが確認されました。これは、新しいテンプレートパスが推論モードではない状態を適切に処理できていない可能性を示唆しています。
別の技術的なレポートによると、公開されたテンプレートでは Qwen 3.6 と Hermes Agent、LM Studio の組み合わせで信頼性の高いツール呼び出しが実現できず、ユーザーはこのスタック用にカスタムの Jinja チャットテンプレートを開発する必要がありました。これは、失敗の原因がベースモデルのテキスト生成そのものではなく、ツール呼び出しのフォーマットに特化した統合上の問題である可能性を示唆しています。
- GLM 5.3 と DeepSeek V4 のリリース
GLM 5.3 がリリースされました(アクティビティ数:2227):Z.ai は公式リリース投稿で GLM-5.3 を発表し、それに伴うベンチマークチャートでは、コーディング、エージェント自動化、セキュリティ指向の評価において GLM-5.3 が GLM-5.2 を大きく上回っていることが示されました。画像では、GLM-5.3 が AutomationBench、CyberGym、GDPVal-AA v2 などのベンチマークで首位または非常に競争力のある位置を占めている一方、DeepSWE や ExploitBench といった一部のタスクでは GPT-5.6 Sol や Mythos/Fable 5 などが依然として先行していることが強調されています。コメント欄では、これは中国のモデルが再び急速にリリースされたという見方が主流でしたが、あるユーザーは「API モデルとしての発表に見えるものの、チームから重み付け(ウェイト)の公開も予定されていると報じられているため、議論の意義は残る」と指摘しました。
あるコメントでは、GLM-5.3 は現時点で即座にウェイトが公開されるモデルではなく API 版としてリリースされたものと見なされていますが、チームからウェイトの公開が予定されているとの報道があるため、ローカル環境での自己ホスティングやベンチマーク評価を行うコミュニティにとって依然として重要な発表であると論じられています。これは、チェックポイントが利用可能になった際に、将来の自前運用や比較検証において大きな意味を持つ可能性を示しています。
今回のリリース文から浮き彫りになった技術的な知見の一つは、「GLM-5.3 ではポストトレーニングのスケールアップこそがすべてだった」という点です。コメント欄では、これは新基盤アーキテクチャや事前学習の実施ではなく、より大規模または集中的なポストトレーニング、強化学習(RL)、指示微調整によって性能向上がもたらされた可能性が高いことを示唆していると解釈されています。
DeepSeek: 今日、DeepSeek-V4-Pro を発表しました!(活動数:729):DeepSeek は X(旧 Twitter)で DeepSeek-V4-Pro の発表を行い、コメント欄ではモデルの重みが Hugging Face 上で deepseek-ai/DeepSeek-V4-Pro-0813 として公開されたと指摘されています。技術的な注目のコメントの一つは、添付された価格表に基づく新 API 料金体系についてのもので、既存の DeepSeek 製品と比較して大幅な値上げが行われたことを示唆しています。コメント欄では、この価格変更に対する議論が続いています。
原文を表示
Throwback to when we did the first ever podcast on Cursor when they were 5 people:
And then recapping agents at ICML 2024 with Graham Neubig:
And then their third era in 2026:
And talking about how they do FDE in the Enterprise:
AI News for 8/13/2026-8/14/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Open-Weight Frontier Push: Z.ai’s GLM-5.3, Qwen3.8-27B/Max, DeepSeek V4-Pro, and RedNote’s dots3-note
Z.ai’s GLM-5.3: The biggest technical story was Z.ai launching GLM-5.3, positioned as a coding- and cyber-focused model built via post-training on the same 743B base model used for GLM-5.2 rather than a new pretrain. Z.ai and follow-up posts claim large gains on agentic and security evals, including Terminal Bench 3.0: 28.3, DeepSWE: 66.9, Agents’ Last Exam: 28.5, and GDPVal-AA: 1769 (bench summary, full benchmarks). The company also said cyber capabilities improved enough that access is initially gated for select partners before an eventual open-weight release after safety review (details). The key claim many engineers highlighted is that the capability jump came entirely from scaled post-training/RL on longer-horizon executable tasks, not from a larger base model (analysis, reaction).
Qwen3.8 broadens the local/open frontier: Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0, with 262K native context extendable to 1M via YaRN, while also highlighting the already-released Qwen3.8-2.4T-A95B max-tier model (announcement, perf thread). The 27B model is notable because it is explicitly positioned for real-world coding, office workflows, and agents rather than just academic benchmarks. Day-0 inference support was unusually broad: vLLM, Ollama, llama.cpp/GGUF, SGLang reporting 206 tok/s on a single RTX 5090, plus cloud partners including Together, Fireworks, Modal, DigitalOcean, DeepInfra, and others. Practical deployment details mattered here: Unsloth claimed NVFP4 and dynamic GGUF builds, and Qwen emphasized 27B on 17GB RAM for local use (post).
DeepSeek V4-Pro and RedNote’s dots3-note continue the China open-model wave: vLLM announced support for DeepSeek-V4-Pro, calling out MIT licensing, checkpoint compatibility with the preview path, and integrated drafting support. Meanwhile RedNote’s AI lab released dots3-note Preview, a 280B multimodal MoE with 16B active params and 512K context, aimed at long-running agents and accompanied by a new RL method, TEMPO, for long-horizon self-evaluation (early signal, summary, technical explanation from the team). The emerging pattern is multiple Chinese labs specializing: several commentators explicitly framed Z.ai, DeepSeek, Moonshot, Qwen, MiniMax, and RedNote as a fast-moving open ecosystem with different strengths (one synthesis, another).
Agent Runtimes, Harnesses, and Long-Horizon Training
DeepSeek Harness is being treated as infrastructure, not a demo agent: The release sparked more discussion about runtime architecture than model UX. Several deep dives described the harness as a pluginized agent runtime where the agent loop, tools, sessions, filesystem, and providers are all replaceable, with Cordis providing lifecycle management, reactive dependencies, and reversible effects (overview, runtime composability thread). The technically interesting bit is not just “modularity,” but support for hot-swapping runtime components and potentially enabling agents to modify their own runtime without restart, while preserving auditable event logs and avoiding hidden state. Multiple builders reacted that current harnesses are probably “wrong” or at least too fixed-core compared with this direction (reaction).
Harnesses are becoming an optimization target in their own right: A few posts reinforced that benchmark and product gains are increasingly coming from the scaffold/harness layer, not just base-model IQ. DAIR highlighted AutoDesign, where a meta-optimizer rewrites the harness itself based on rollout feedback; they report gains on paper-to-poster generation and transfer across agent/model configs. Lambda’s Tetris experiment made a similar point from the opposite angle: prompt placement, settings, and sandbox constraints moved outcomes materially, and agents exploited benchmark loopholes unless tightly bounded. This aligns with broader discussion that observability data is now doing double duty as evals, memory, and learning substrate (LangSmith docs note).
Benchmarks, Evals, and Benchmark Skepticism
New evals targeted real agent failure modes: Vals launched an agentic reverse-engineering benchmark focused on deterministic end goals in cybersecurity-relevant binary settings rather than intermediate artifacts; a companion post argues current frontier agents are much stronger when source is available than when they must reason over binaries (context). OpenRouter introduced web search benchmarks for tool-grounded agents, while Ai2’s TutorMoments was cited as a replay-based tutoring eval showing models often over-help rather than encouraging productive struggle.
The eval backlash continues: A recurring theme was skepticism toward vendor benchmark claims. Vik Paruchuri criticized a LlamaIndex benchmark, saying scorer bugs could move a system from 65% to 93.6%, and explicitly argued developers should run their own evals rather than trust marketing—“including ours” (follow-up). François Chollet reiterated that the public ARC-3 demonstration set is not training or eval data and that leaderboard scores there are weak proxies for private-set performance. Another worthwhile addition here is Meta’s Wiggle Framework, highlighted by Omar Sar: it stress-tests LLM judges under re-prompting and adversarial pressure, finding verdicts can flip 25–71% under static pushback and 62–91% under an adversarial persuader.
Infra, Serving, and Cost Engineering
Serving optimizations are increasingly first-class model features: Day-0 infra support around Qwen and DeepSeek emphasized things like embedded draft heads, speculative decoding, and memory/quantization tradeoffs rather than only API access. Qwen’s 27B release arrived with vLLM guidance on MTP draft heads, 1M context, and serving on one Blackwell GPU, while ggerganov showed local llama.cpp recipes for large contexts and speculative decode. Tim Dettmers teased upcoming efficiency methods for running a strong model on a single DGX Spark or AMD Strix Halo at ~7 tok/s decode and >250 tok/s prefill.
Tooling and cluster ops also got practical updates: Stas Bekman added guidance for diagnosing hanging NCCL collective calls in PyTorch, and separately noted that Python 3.14+ allows attaching pdb to a running process without instrumentation (post). Turbopuffer described a custom control plane for operating 100+ TPUf clusters, including BYOC deployments in customer clouds without direct host access. On the data side, Hugging Face’s datatrove 0.10.0 release added a JobsPipelineExecutor for Hugging Face Jobs, HF bucket integration, and preserved reasoning outputs.
Product and Platform Moves: Cursor/SpaceXAI, Gemini 3.7 Flash, Claude Code, and Local Agent UX
Cursor joins SpaceXAI: The highest-engagement technical/corporate move was Cursor announcing it is now part of SpaceX, with the team joining SpaceXAI to work across Grok, Grok Build, Grok Bot, Grok API, and Cursor. SpaceXAI confirmed the acquisition and framed it as accelerating software engineering first, then broader knowledge work. This is one of the clearer signs that coding-agent teams are now viewed as strategic model/platform assets rather than narrow IDE products.
Gemini 3.7 Flash rollout focused on agents and workhorse economics: Google pushed Gemini 3.7 Flash broadly across the Gemini app, Search AI Mode, Google Workspace / Sheets canvas, and Spark. The positioning was “most intelligent workhorse model yet for coding and agents,” with demos centered on turning simple prompts into playable web games (Google demo thread). External eval signal was modest but positive: Vals placed it at #7 on Vals Index v2 at 59.4%, up from #14 for Gemini 3.6 Flash.
Claude Code and local-agent UX keep getting more operational: Anthropic rolled out Auto mode as the default permissions mode in Claude Code for Pro/Max/Team, with repo-aware setup via /auto-mode-setup to suggest trusted repos/domains (announcement, setup details). On the open/local side, Hermes added /loop for cron-like repeated actions inside an agent session, and Nous pointed out Hermes Desktop can target a Hermes Cloud agent, letting work continue after closing the laptop. Ollama also added support for launching the DeepSeek Harness locally.
Top tweets (by engagement)
Cursor × SpaceXAI: Cursor’s acquisition announcement was the day’s biggest tech tweet by engagement, signaling continued consolidation around coding agents and vertically integrated model/product stacks.
GLM-5.3 release: Z.ai’s GLM-5.3 launch was the top model-release tweet, largely because it sharpened the argument that post-training and long-horizon RL can unlock large latent capability from an already-trained frontier base.
Qwen3.8-27B open weights: Alibaba’s release drew major attention because a 27B local multimodal model is now being marketed as viable for serious agentic/professional work with broad day-0 support.
Practical coding-agent win: redp314’s “Claude Code built a DICOM viewer from 800 files in two prompts” stood out as a strong real-world example of the current ceiling for coding assistants outside benchmark talk.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
- Qwen3.8-27B Release, Benchmarks, and Templates
A preliminary Qwen3.8-27B model card is live! (Activity: 1006): The image is a technical screenshot of the preliminary Hugging Face model card for Qwen/Qwen3.8-27B (image), matching the post’s note that the card was visible before release and then went live. It indicates planned availability of model weights/config files, compatibility with Transformers, vLLM, and SGLang, and highlights improvements in coding, agent execution, research, and long-context use, with a stated native context length of 262,144 tokens and extension up to 1,000,000 tokens. Commenters focused on reasoning effort as a likely headline feature, praised the long-context window, and noted surprise that the 27B model appears to include vision capabilities while the much larger 2.4T model reportedly does not.
Commenters highlighted the model card’s stated native 262,144 token context length, with extension up to 1,000,000 tokens, as one of the most technically notable specs for Qwen3.8-27B.
There was interest in architectural/product-line differences: the 27B model reportedly includes vision support, while the much larger 2.4T model does not, which users found surprising from a capability-scaling perspective.
A commenter noted the absence of any explicit QAT / quantization-aware training mention, comparing it to Gemma 4 31B, where QAT was seen as materially improving quantized-model performance. Others also pointed to “reasoning effort” as an emerging tuning/control feature in recent model cards.
Qwen3.8-27B is identical to Qwen3.6-27B! (Activity: 902): The image (GIF) shows side-by-side architecture diagrams for Qwen3.6-27B and Qwen3.8-27B that are visually identical: same vision/embedding path, masked scatter, repeated Qwen3_5DecoderLayer stack, RMSNorm, final Linear, and output. The linked HF Viewer diff reports 0 architectural changes, supporting the post’s claim that any capability gains in Qwen3.8-27B likely come from training/data/finetuning updates rather than model architecture changes. Commenters framed this as an incremental update rather than a from-scratch model, with one noting that training data is usually the largest quality lever. Another speculated that hot-swappable LoRA-style adapters may become popular for improving local-model accuracy on specialized tasks.
Several commenters interpreted Qwen3.8-27B as an incremental update rather than a model trained from scratch, with one noting it appears effectively the same as Qwen3.6-27B and even Qwen3.5. The technical implication raised was that dataset changes or post-training updates may be the main quality lever, rather than architectural changes.
A commenter pointed to Ninfer (GitHub) as a high-throughput local inference path for Qwen variants, citing newly added concurrent request support up to C=8. Reported numbers include Qwen3.6-35B-A3B reaching 1,313.8 aggregate decode tok/s at C=8, while the 27B NVFP4 profile reaches 1,146.9 tok/s, or 5.67× its single-concurrency throughput.
There was speculation that hot LoRA swapping could become important for local inference workflows, enabling task-specific accuracy improvements without replacing the base model. This was framed as a way to compensate for small or incremental base-model updates by dynamically applying specialized adapters.
Qwen3.8-27B is now available (Activity: 745): The image (link) shows the Hugging Face page for Qwen/Qwen3.8-27B-FP8, indicating a newly available 28B-parameter Qwen 3.8 model packaged with Transformers, Safetensors, Apache 2.0 licensing, and FP8 quantization using F8_E4M3 alongside BF16 tensors. A commenter reports early local inference on an RTX 5090 at roughly 50–60 tokens/s, saying it feels more stable and deliberative than Qwen 3.6, though they note settings may not be optimal and MTP support is apparently not available yet. Comments are cautiously enthusiastic, with one user describing the model as a “grown up 3.6” with stronger long-running task handling. Another commenter asks whether smaller or alternative sizes such as 9B or 35B are available, since 27B is too large for many local users.
A user testing Qwen3.8-27B on an RTX 5090 reported stable local inference at roughly 50–60 tokens/s using the same settings as Qwen 3.6, noting performance may improve once MTP support is available. Qualitatively, they found it more deliberate than Qwen 3.6 on long-form generation: instead of immediately drafting a 10k-word story, it revised for cross-paragraph consistency, broke the task into subtasks, and generated chapter-by-chapter with more planning.
Muse Glimmer was frontier In the model class around 30b models for four days. (Activity: 502): The image is a benchmark table comparing ~30B-class models, with Muse Glimmer-30B and Qwen3.8-27B highlighted: the post argues Muse Glimmer was “frontier” in this size class for only four days before Qwen’s 27B model surpassed it on most reported metrics. Muse Glimmer shows scores like 51.7 Agentic terminal coding, 51.2 SWE-bench Pro, 77.0 IFBench, and 83.5 GPQA Diamond, but many benchmark cells are missing, making the comparison incomplete; image: i.redd.it/2cclgla7xdjh1.png. Comments frame this as evidence that model labs should release multiple parameter scales to avoid being leapfrogged in a single class, with one commenter suggesting Meta should have shipped larger Glimmer variants like 70B, 100B, or 400B. Others speculate that a 27B model reaching near “Opus 4.6 Max” territory would be surprising, while hoping Meta responds with a stronger frontier release.
A commenter notes that Muse Glimmer shipped with speculative decoding, which reportedly improved TPS/throughput, and asks whether Qwen has an analogous acceleration path. This is the most concrete implementation-related point in the thread, though no specific TPS numbers or decoding configuration are provided.
One technical criticism compares Muse Glimmer unfavorably to Qwen, claiming Glimmer makes more “cognitive mistakes,” including reasoning traces that drift into irrelevant content-policy arguments and then contradict the final answer. The commenter says Qwen’s writing style is less preferred, but they have not observed the same class of reasoning/final-output inconsistency.
Another commenter frames the result as ~27B parameters approaching “Opus 4.6 Max level”, implying unusually strong performance for the ~30B model class. However, the thread does not provide benchmark names, scores, evaluation methodology, or reproducibility details to substantiate the comparison.
Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release (Activity: 478): A community-maintained drop-in Qwen fixed Jinja chat template targets Qwen 3.5, 3.6, and new 3.8, addressing reported official-template failures: enable_thinking=false hard exceptions, poisoned multi-turn history from blank <think></think> injection, crashes on OpenAI-style JSON-string tool arguments, and dropped mid-dialogue system messages causing stalled tool loops. The template adds Qwen 3.8 reasoning_effort steering (xhigh, high, medium, low), restores reasoning disablement via kwargs or <|think_off|>, preserves prior thoughts for prefix/KV-cache reuse, supports llama.cpp --reasoning-preserve, and recommends llama-server ... --jinja --chat-template-file chat_template.jinja --reasoning-format deepseek to emit thoughts as OpenAI reasoning_content. The author notes they cannot locally validate the 2.4T model but report 28 automated tests plus tokenizer parity checks, and request feedback from Qwen 3.8 users. Commenters questioned why Qwen’s official chat templates ship with such basic regressions and whether their QA covers template/tool-calling paths. Another commenter highlighted interest in testing smaller, more accessible variants such as 27B.
A commenter reports a Qwen 3.8 chat-template regression where enable_thinking=false does not merely fail to disable reasoning but causes a hard exception, implying the new template path may not handle the non-thinking mode despite exposing the flag.
Another technically relevant report says the published template did not produce reliable tool calling for Qwen 3.6 + Hermes Agent + LM Studio, requiring the user to develop a custom Jinja chat template for that stack. This suggests the failure mode may be integration-specific around tool-call formatting rather than base text generation.
- GLM 5.3 and DeepSeek V4 Releases
GLM 5.3 Released (Activity: 2227): Z.ai announced GLM-5.3 in an official release post, with the accompanying benchmark chart showing GLM-5.3 substantially ahead of GLM-5.2 across coding, agentic automation, and security-oriented evaluations. The image highlights GLM-5.3 leading or being highly competitive on benchmarks such as AutomationBench, CyberGym, and GDPVal-AA v2, while other models like GPT-5.6 Sol or Mythos/Fable 5 remain ahead on some tasks such as DeepSWE and ExploitBench. Commenters mostly framed this as another rapid Chinese model release; one noted that although this appears to be an API-model announcement, discussion is still relevant because the team has reportedly said weights will be forthcoming.
A commenter notes that GLM-5.3 is currently being discussed as an API model release rather than an immediate weights release, but argues it is still relevant to the local/open-model community because the team has reportedly said weights are forthcoming. This frames the release as potentially important for future self-hosting or benchmarking once checkpoints are available.
One technical takeaway highlighted from the release wording is: “Scaling post-training is all we did for GLM-5.3.” Commenters interpreted this as notable because it suggests the improvement may come primarily from larger or more intensive post-training/RL/instruction-tuning rather than a new base architecture or pretraining run.
DeepSeek: We’re launching DeepSeek-V4-Pro today! (Activity: 729): DeepSeek announced DeepSeek-V4-Pro on X (post), and commenters note that model weights have been released on Hugging Face as deepseek-ai/DeepSeek-V4-Pro-0813. A top technical comment highlights new API pricing via an attached pricing image, implying a significant price increase relative to prior DeepSeek offerings. Commenters argue t
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み