今日は何も大きな出来事はありませんでした
騰訊が Apache 2.0 ライセンスで公開した大規模 MoE モデル「Hy3」は、推論最適化と高い性能によりオープンソース界隈に大きな衝撃を与え、業界の最上位層への参入を確立しました。
キーポイント
高性能なオープンモデル Hy3 の公開
騰訊が Apache 2.0 で公開した 295B MoE モデル「Hy3」は、21B アクティブパラメータと 256K コンテキスト長を備え、推論・コーディング・エージェントタスクで競合他社を上回る性能を示しています。
リリース直後の成熟した推論サポート
vLLM がローンチ時から Hy3 をネイティブサポートし、NVIDIA および AMD 環境での動作を検証済みです。特に MTP による推測デコーディングや FP8 モードの統合により、レイテンシが大幅に削減されています。
業界内での評価と比較
Hy3 は GLM-5.2 と比較され、ベンチマーク結果次第ではオープンソースラボの最上位層に騰訊を加える可能性があると議論されており、コミュニティからの反応も非常に活発です。
重要な引用
Tencent released Hy3 under Apache 2.0, a 295B MoE with 21B active parameters
Hy3 runs natively in vLLM from launch with tool-call and reasoning parsers
reported gains of up to 2.95x on mixed-length decode
影響分析・編集コメントを表示
影響分析
本記事は、大規模モデルの性能競争において「オープンソース」の壁がさらに高まりつつあることを示しています。特に推論エンジン(vLLM)との深い統合により、単にパラメータ数が多いだけでなく、実運用レベルでの効率性も確保された点が業界全体のパラダイムシフトを加速させる要因となります。
編集コメント
「静かな日」というタイトルとは裏腹に、Hy3 のリリースはオープンソース LLM エコシステムにおける重要なマイルストーンです。推論効率の劇的な向上が、大規模モデルの実用化をさらに加速させるでしょう。
静かな一日。
2026年7月4日〜7月6日のAIニュース。12のサブレッド、544 の Twitter、およびさらに多くのDiscordをチェックしました。AINews のウェブサイトでは過去のすべての号を検索できます。念のため、AINews は現在 Latent Space のセクションの一部です。メールの頻度を選択的に設定(購読または解除)することができます!
AI Twitter リキャップ
Tencent Hunyuan の Hy3 リリースとオープンウェイトの最前線
- Hy3 は本格的なオープンモデルとして登場:Tencent は Apache 2.0 ライセンスの下で Hy3 をリリースしました。これは 295B モデル(MoE)で、アクティブパラメータは 21B、エクスパート数は 192、トップ-8 ルーティング、GQA(Grouped Query Attention)、コンテキスト長は 256K、推論デコーディング用の 3.8B MTP レイヤーを備えています。複数の投稿では、推論、コーディング、エージェントタスクにおいてより大規模なシステムと競合可能であると位置づけられ、特にツール呼び出しの安定性や幻覚(hallucination)防止などの信頼性の向上に重点が置かれています @eliebakouch, @HuggingPapers, @ShunyuYao12。
- 推論サポートは例外的に日付 0 で成熟していました:@vllm_project は、Hy3 がローンチ時から vLLM でネイティブに実行可能であり、ツール呼び出しおよび推論パーサー、MTP(Multi-Token Prediction)によるスペキュレーティブデコーディング、NVIDIA および AMD での検証済みサポートを備えていると述べています。続報では、Tencent の生産用カーネルが vLLM メインブランチにアップストリームされたことが詳細に説明され、負荷分散型デコードスケジューリングや融合 FP8 MoE(Mixture of Experts)サービングが含まれており、デフォルトバックエンドと比較して混合長デコードで最大 2.95 倍の向上、TTFT(Time To First Token)で約 24%、TPOT(Time Per Output Token)で約 17% のレイテンシ削減が報告されています。コミュニティからの反応は強く、@Teknium はすぐに Nous Portal で Hy3 を 2 週間無料公開しました。
- より広範なオープンモデルの文脈:Hy3 は直ちに GLM-5.2 と比較され、一部の投稿者はベンチマークとバイブテストの結果が裏付けられるなら、Tencent が今やオープンソースラボの最上位階層に加わったと主張し @teortaxesTex、一方で他の人々は依然として GLM-5.2 を実用上で現在利用可能な最高のオープンウェイトモデルであると維持しています @tinygrad, @mbusigin。全体の結論:オープンのフロンティアは急速に圧縮されており、競争は単なるリーダーボードの差分ではなく、デプロイメントの堅牢性 increasingly についてのものである。
エージェントベンチマーク、ハルネス、および長期実行メモリ
- AutomationBench-AA はより現実的なエージェント評価を追加しました:@ArtificialAnlys が Zapier の AutomationBench に対して独立したリーダーボードを立ち上げ、657 のタスクと 40 のシミュレートされた SaaS アプリ(SaaS: Software as a Service)において、目的達成とガードレール(安全・制約ルール)の両方を考慮してエージェントを評価しました。Claude Fable 5 が 48.6% で首位に立ち、Opus 4.8 の 48.5% をわずかに上回りました。Gemini 3.5 Flash は 42.6%、GPT-5.5 xhigh は 42.1% です。ランキング自体よりも興味深いのは、すべてのモデルが依然としてビジネスルールを破っている点であり、特に Gemini は「目的達成ごとのガードレール違反」およびコスト効率において顕著に強みを示しました。オープンウェイト(Open weights: 公開された重みパラメータを持つモデル)は依然として明確に後れを取っており、リストアップされたオープンモデルの中で最も良かったのは GLM-5.2 max の 27.8% です。
- キャパビリティ指数が多次元化しています:Artificial Analysis はまた、単一のスカラー値によるモデルスコアから脱却するため、6 つのドメイン固有の指数を導入しました。これらは「財務・会計」「法務」「医療・ヘルスケア」「戦略・運営」「エンジニアリング」「経済学」です(@ArtificialAnlys)。見出しは慣れ親しんだ内容—Claude Fable 5 と Opus 4.8 のフォールバックが首位—でしたが、より有用な洞察は、ドメインによってランキングがどのように劇的に入れ替わるか、そしてパフォーマンスと価格のフロンティアがいかに急峻になっているかという点です。これは @fchollet の主張とも一致しており、「タスクあたりのコストを伴わないベンチマークスコアの報告は、もはや意味をなさない」という立場です。
- メモリと検索は、永続型エージェントにおけるボトルネックのままです:この分野で注目された論文が 2 つあります。まず、A-TMA は「ゴーストメモリ」、つまり長期稼働するアシスタントにおいて古く現在の事実が混在して検索される問題を解決します。LTP ベンチマークでは、Graphiti に A-TMA を追加することで競合の精度が絶対値で +0.240 向上したと報告されています @omarsar0。次に、ReContext はトレーニング不要の長文コンテキスト推論用ハーンスであり、回答生成直前にモデル内部のエビデンスを再生することで、8 つの 128K データセット全体でエビデンスの利用効率を向上させます @dair_ai。百万トークン単位のインコンテキスト検索である BlockSearch と組み合わせることで(@dair_ai)、テーマは明確です:より優れたメモリ動作は、トレーニング時だけでなく推論時にエンジニアリングされるようになりつつあります。
Anthropic の J-Space / グローバル・ワークスペースに関する結果
- 機械的解釈可能性が中心に据えられました:Anthropic は、Claude にグローバル・ワークスペースのような内部構造が存在すると主張する研究を発表しました。これは彼らが「J-space」と呼ぶ活性化の小さなサブセットを中心としたものです @AnthropicAI, @AnthropicAI。核心的な主張は思考連鎖(chain-of-thought)の抽出ではなく、報告、調整、柔軟な推論のために利用可能な特権的な内部表現基盤の特定です。また、Anthropic はオープンウェイトモデル向けの Neuronpedia デモも提供しました @AnthropicAI。
- なぜ研究者たちが関心を持ったか:解釈可能性の研究者たちは、この結果を以前の公開研究よりもモデルの「作業記憶」または内部ワークスペースに対するより強力な証拠とみなしました。見方については異論があってもです。@NeelNanda5 はこれを作業記憶に似たメカニズムに関するこれまでに得られた最良の証拠と呼びました。@Jack_W_Lindsey は、この特権的な空間を理解することが LLM(大規模言語モデル)の認知における鍵となる可能性があると主張しました。また、投稿では実用的な安全性の観点も強調されました:このワークスペースは、隠れた概念を表面化させたり、プロンプトインジェクションを検出したり、内部でのサボタージュに関連する特徴が言葉として表出される前に暴露したりできると報告されています @mlpowered, @LiorOnAI, @omarsar0。
- しかし「意識」という言語は議論の対象となりました:Anthropic の公開による枠組みには強い反発がありました。支持者たちは、この結果は現象的意識ではなくアクセスの機能アナログを示唆していると述べました @BorisMPower。一方、批判者たちは、企業が特権的な潜在活性化と意識を混同することで主張しすぎていると指摘しました @AlanCowen。共感的な見解を持つ人々でさえも、より大きな物語は哲学ではなく、モデルの監査や誘導のための新たな介入点であることに焦点を当てるべきだと強調しました。
推論、サービング、およびシステム効率
- Speculative decoding は引き続き注目されるインフラ技術です:@lmsysorg が SGLang に DSpark を追加し、信頼度に基づいた可変長の検証を実現しました。その狙いは、高負荷時にすべてのドラフトトークンを検証するのではなく、固定予算型の推測手法に比べてスループットとレイテンシのトレードオフを改善することです。DeepSeek-V4-Pro は B300 上でバッチサイズ=1 の条件下で 383.7 tok/s を達成したとの報告があります。また Microsoft も、@code および @pierceboggan が発表後、GPT-5.5 のレイテンシとトークン効率を向上させるため GitHub Copilot ハーネス内でのプロンプトレベルの最適化について議論しました。
- 推論効率が戦略的なボトルネックとなりつつあります:@jon_durbin は、トレーニングだけでなく推論そのものが「ゲーム全体」であると主張し、すべてのデータパイプライン、RL ループ、エージェントランタイムが最終的にテスト時の計算リソースとして結実するためです。この視点はより低レベルなカーネルワークにも表れており、Chutes が MiniMax MSA および GatedDeltaNet-2 において大幅な高速化を報告しました。これには RTX Pro 6000 / SM120 上でのスパースアテンショントレーニングの約 7 倍の改善や、統合された FP8 カーネルの向上が含まれます @jon_durbin。
- モデルサービングを超えたインフラリリース:Cloudflare は Workers Cache を発表しました。これは標準的な HTTP ヘッダーを介して設定可能な Worker エントリーポイント前に配置される地域階層型キャッシュです @Cloudflare。OpenAI は GPT-Realtime-2.1-mini をリリースし、推論機能とツール使用機能をミニリアルタイムラインに導入すると同時に、前世代のミニモデルと同じ価格で提供しました。またキャッシュ改善により p95 レイテンシが 25% 以上削減されたとの主張があります @OpenAIDevs, @OpenAIDevs。
World Models, Speech, and Document AI
- MIRA は注目すべき世界モデルのデモです:General Intuition と Kyutai、そして Epic Games が共同で、1 万時間のボット収集データを用いてトレーニングされた Rocket League 向けのプレイ可能なマルチプレイヤー世界モデル「MIRA」を発表しました。これは @gen_intuition によって紹介されました。同モデルはリアルタイムで 20fps で動作し、@TheRundownAI による投稿では、明示的な物理エンジンやレンダリングエンジンを使用せずに、単一の NVIDIA B200 グラフィックボード上で 2v2 の試合全体を処理する 50 億パラメータモデルのデモが紹介されました。これは、動画・世界モデルの研究が「おもちゃのようなデモ」から「インタラクティブなシミュレーター」へと移行していることを示す最も明確な信号の一つでした。
- 音声技術分野は依然として激しい競争状態にあります:AssemblyAI は、AA-WER Streaming および文脈による事前提示(コンテキストプリミング)に対応し、通話中に再接続せずに更新可能なストリーミング型音声認識(STT: Speech-to-Text)モデル「Universal-3.5 Pro Realtime」をリリースしました。このモデルは 4.1% の WER を達成しており、@ArtificialAnlys が紹介しています。テキスト読み上げ(TTS: Text-to-Speech)分野では、Artificial Analysis によると、「Speechify Simba 3.2」が現在、Speech Arena で 1233 Elo のスコアを記録し、Gemini 3.1 Flash TTS、Sonic 3.5、Inworld Realtime TTS 1.5 Max を引き離して首位に立っています。また、上位モデルの中で最も安価な選択肢でもあります(@ArtificialAnlys)。
- ドキュメントコンテキストパイプラインはデフォルトでマルチモーダル化されつつあります:LlamaIndex と LanceDB は、不揃いな PDF 向けの検索パイプラインを公開しました。このパイプラインではページを分離し、チャンク(断片)化し、抽出されたアセットをリンクされたマルチモーダルテーブルに格納することで、ラベル付き ESG レポートベンチマークにおいて「任意のページヒット率@5」で 82%、「回答精度」で 74% を達成しました。これは、エージェント向けの専用「ドキュメントコンテキストレイヤー」を主張する Jerry Liu のより広範な議論と相まっており(@jerryjliu0)、両者は補完関係にあります。
エンゲージメント上位のツイート
- Anthropic のグローバルワークスペースに関する論文がエンゲージメントを支配し、Claude の内部ワークスペース/J-スペースに関する主要な発表は、@AnthropicAI における他のすべての話題を大きく上回りました。
- Tencent Hy3 は、オープンソースの競争力や展開について議論する技術系アカウントを中心に、純粋なモデルリリースとしての最大のストーリーとなりました @teortaxesTex, @ShunyuYao12。
- MIRA のプレイ可能なワールドモデルは、@gen_intuition による注目すべきマルチモーダル/システムデモでした。
- Will Depue の「データのためのスターゲイト」スレッドは、最も実質的な戦略投稿であり、データ収集が計算リソース単独ではなく、フロンティアラボにおけるボトルネックとなる制約かつ潜在的な参入障壁(モート)になると主張しました @willdepue。
- John Carmack のメモリシステムに関するスレッドは、推論ハードウェアが大規模モデルのサービス提供において HBM よりもはるかに安価なメモリ階層と、決定論的なアクセスパターンを活用できると主張することで、大きな技術的関心を集めました @ID_AA_Carmack。
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Large Open-Weight MoE Model Releases
- longcat 2.0 (1.6T, ~48B active) の重みは、MIT ライセンスの下で公開されました(アクティビティ:638): elie と ModelScope からの発表および LongCat 2.0 のブログ記事における技術詳細により、LongCat 2.0 の重みが MIT ライセンスの下で公開されました。このモデルは非常に大規模な MoE システムであり、総パラメータ数は 1.6T、推論時のアクティブ・パラメータ数は約 48B です。コメント投稿者たちは、公開された重みが BF16 で約 3.55TB、FP8 で約 2.05TB を占有することを指摘し、その巨大なサイズによる実用的な展開の負担を強調しました。また、中国版グーグル/Uber Eats と例えられる美团(Meituan)が、完全に国内製の中国製チップでこのモデルを訓練したと報じられていることから、地政学的・市場的な意義について議論が巻き起こっています。
コメント投稿者たちは、LongCat 2.0 の規模と展開に必要なリソースに焦点を当てました。総パラメータ数は 1.6T で、アクティブ・パラメータは約 48B とされており、これはスパース/MoE 型アーキテクチャを示唆しています。あるユーザーは、公開された重みが BF16 で約 3.55TB、FP8 で約 2.05TB を必要とすることを指摘し、ローカルストレージや推論インフラストラクチャを計画する者にとって重要な情報であると述べています。
- 技術的な観点として、美团(Meituan)がモデルを 100% 国内製の中国製チップで訓練したという報じられた事実が取り上げられ、コメント投稿者たちはこれを AI ハードウェアサプライチェーンの自立において重要であるとの文脈で捉えました。これは、従来の AI ラボではなく、グーグルと Uber Eats を組み合わせたような主要な中国インターネット企業としての美团の役割を考えると、特に注目すべき点です。
- 複数のユーザーが、寛容な MIT ライセンスに注目し、Qwen や DeepSeek といった最先端のオープンモデルとのベンチマーク計画を打ち出しました。総パラメータ数 1.6T、アクティブパラメータ数は約 48B に過ぎず、かつ重み(weights)が公開されているという組み合わせは、推論ツールがこのアーキテクチャを効率的にサポートする限り、他の高機能な MoE(Mixture of Experts:専門家混合モデル)オープンモデルと比較対象として実用的である可能性を示唆しています。
- Tencent Hy からの新しいオープンモデル「Hy3」(総パラメータ数 295B、アクティブパラメータ数 21B - Apache 2.0 ライセンス)(活動状況:604): Tencent は Hugging Face にプレビュー版ではない Hy3 モデルコレクションをリリースしました。これは 295B パラメータの MoE でアクティブパラメータ数は 21B とされ、以前は制限的なコミュニティライセンスでしたが、現在は Apache 2.0 ライセンスの下にあります。コメント投稿者はリンクされたベンチマークチャートを指摘し、「Hy3-Preview に対して非常に印象的な claimed gains(主張される性能向上)が示されており」、実世界の性能が報告結果と一致すれば、高機能なローカル/家庭用推論環境でも有用になる可能性があると述べています。主な議論はライセンス変更について肯定的であり、コメント投稿者は地理的・制限的なライセンスから Apache 2.0 へ移行することを最も重要な改善点と捉えており、特に Tencent が最近 Apache ライセンスで公開した翻訳モデル群と併せて注目されています。
コメント投稿者たちは、Tencent Hunyuan Hy3 が総パラメータ数 295B でアクティブパラメータが 21B の大規模 MoE(Mixture of Experts)モデルとしてリリースされた点を指摘しています。あるユーザーは、HY3-Preview に対するベンチマークでの向上幅が実務ワークロードでも有効に転用されるならば、「ハイエンドな家庭用セットアップ」において関連性を持つほど十分であると述べています。ローカル展開が可能かどうかを決定づけるものとして、特に GGUF(GGML Unified Format)形式の量子化モデルの実用推論利用への関心が寄せられています。
- 技術的に重要なライセンス変更が指摘されました:Tencent は、韓国、英国、EU などの地域での使用を制限していたより restrictive な「コミュニティ」ライセンスから、Apache 2.0 ライセンスへ移行したと報じられています。コメント投稿者たちはこれを重要視しており、これは Tencent の最近の Apache ライセンス付き翻訳モデルとも整合性があり、より広範な商用利用や研究再利用を可能にするためです。
- ある投稿者は Hy3 を Qwen や MiniMax に対する潜在的な代替案として位置づけ、現在の主要なオープンウェイト中国製モデルファミリーとベンチマークおよび実世界でのパフォーマンスで競合できるかどうかに関心があることを示唆しています。
新モデル:GigaChat3.5-432B-A28B(当日 GGUF サポート付き!) (アクティビティ:439): Sberbank/ai-sage が Huggi で GigaChat3.5-432B-A28B をリリースしました
原文を表示
a quiet day.
AI News for 7/04/2026-7/06/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews' website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Tencent Hunyuan’s Hy3 Release and the Open-Weight Frontier
- Hy3 lands as a serious open model: Tencent released Hy3 under Apache 2.0, a 295B MoE with 21B active parameters, 192 experts / top-8 routing, GQA, 256K context, and a 3.8B MTP layer for speculative decoding. Multiple posts framed it as competitive with much larger systems on reasoning, coding, and agentic tasks, with particular emphasis on reliability improvements like tool-calling stability and anti-hallucination work @eliebakouch, @HuggingPapers, @ShunyuYao12.
- Inference support was unusually day-0 mature: @vllm_project said Hy3 runs natively in vLLM from launch with tool-call and reasoning parsers, MTP speculative decoding, and validated support on NVIDIA and AMD. A follow-up detailed Tencent production kernels now upstreamed into vLLM main, including load-balanced decode scheduling and fused FP8 MoE serving, with reported gains of up to 2.95x on mixed-length decode and latency reductions of roughly 24% TTFT and 17% TPOT versus default backends @vllm_project. Community reaction was strong enough that @Teknium quickly made Hy3 free on Nous Portal for two weeks.
- Broader open-model context: Hy3 was immediately compared against GLM-5.2, with some posters arguing Tencent has now joined the very top tier of open-source labs if the benchmark and vibe-test results hold @teortaxesTex, while others still maintained GLM-5.2 as the best currently usable open-weight model in practice @tinygrad, @mbusigin. The net takeaway: the open frontier is compressing fast, and the competition is increasingly about deployment robustness rather than just raw leaderboard deltas.
Agent Benchmarks, Harnesses, and Long-Running Memory
- AutomationBench-AA adds a more realistic agent eval: @ArtificialAnlys launched an independent leaderboard for Zapier’s AutomationBench, evaluating agents across 657 tasks and 40 simulated SaaS apps with both objectives and guardrails. Claude Fable 5 led at 48.6%, narrowly ahead of Opus 4.8 at 48.5%, with Gemini 3.5 Flash at 42.6% and GPT-5.5 xhigh at 42.1%. More interesting than the ranking: every model still breaks business rules, and Gemini looked notably strong on objective-per-guardrail-violation and cost efficiency. Open weights remain meaningfully behind, with GLM-5.2 max the best listed open model at 27.8%.
- Capability indices are becoming multidimensional: Artificial Analysis also introduced six domain-specific indices—Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, Economics—to move past single scalar model scores @ArtificialAnlys. The headline was familiar—Claude Fable 5 plus Opus 4.8 fallback leads—but the more useful insight is how sharply rankings reshuffle by domain and how steep the price/performance frontier has become. This aligns with @fchollet, who argued that reporting benchmark scores without cost per task is increasingly meaningless.
- Memory and retrieval remain bottlenecks for persistent agents: Two papers got traction here. First, A-TMA tackles “ghost memory,” where stale and current facts are retrieved together in long-running assistants; on the LTP benchmark, adding it to Graphiti reportedly improves conflict accuracy by +0.240 absolute @omarsar0. Second, ReContext is a training-free long-context inference harness that replays model-internal evidence right before answer generation, improving evidence utilization across eight 128K datasets @dair_ai. Combined with BlockSearch for million-token in-context retrieval @dair_ai, the theme is clear: better memory behavior is increasingly being engineered at inference time, not just trained in.
Anthropic’s J-Space / Global Workspace Results
- Mechanistic interpretability took center stage: Anthropic released research claiming a global-workspace-like internal structure in Claude, centered on a small subset of activations they call J-space @AnthropicAI, @AnthropicAI. The core claim is not chain-of-thought extraction, but identification of a privileged internal representational substrate that appears available for report, modulation, and flexible reasoning. Anthropic also shipped a Neuronpedia demo for open-weight models @AnthropicAI.
- Why researchers cared: Interpretability researchers treated this as stronger evidence for a model “working memory” or internal workspace than prior public work, even if they disagreed with the framing. @NeelNanda5 called it the best evidence yet for a working-memory-like mechanism. @Jack_W_Lindsey argued understanding this privileged space could be key to LLM cognition. Posts also highlighted practical safety angles: the workspace can reportedly surface hidden concepts, detect prompt injections, and expose internal sabotage-related features before they are verbalized @mlpowered, @LiorOnAI, @omarsar0.
- But the “consciousness” language was contested: Anthropic’s public framing invited strong pushback. Supporters said the results suggest a functional analog of access consciousness rather than phenomenal consciousness @BorisMPower, while critics argued the company was overclaiming by conflating privileged latent activation with consciousness @AlanCowen. Even some sympathetic takes emphasized the bigger story is a new intervention point for auditing and steering models, not philosophy.
Inference, Serving, and Systems Efficiency
- Speculative decoding remains hot infrastructure: @lmsysorg added DSpark to SGLang for confidence-driven, variable-length verification. The pitch is that under high load it avoids verifying every draft token, improving the throughput/latency tradeoff relative to fixed-budget speculative methods; DeepSeek-V4-Pro reportedly reached 383.7 tok/s at batch=1 on B300. Microsoft also discussed prompt-level optimization of GPT-5.5 in the GitHub Copilot harness to improve latency and token efficiency after launch @code, @pierceboggan.
- Inference efficiency is increasingly the strategic bottleneck: @jon_durbin argued that inference, not training alone, is now “the whole game,” because every data pipeline, RL loop, and agent runtime ultimately cashes out as test-time compute. That perspective also showed up in lower-level kernel work: Chutes reported major speedups for MiniMax MSA and GatedDeltaNet-2, including ~7x sparse-attention training improvements on RTX Pro 6000 / SM120 and better fused FP8 kernels @jon_durbin.
- Infra releases beyond model serving: Cloudflare launched Workers Cache, a regionally tiered cache in front of Worker entrypoints configured via standard HTTP headers @Cloudflare. OpenAI shipped GPT-Realtime-2.1-mini, bringing reasoning and tool use to the mini realtime line at the same price as the prior mini, alongside claimed 25%+ p95 latency reductions from caching improvements @OpenAIDevs, @OpenAIDevs.
World Models, Speech, and Document AI
- MIRA is a notable world-model demo: General Intuition and Kyutai, with Epic Games, introduced MIRA, a playable multiplayer world model for Rocket League trained on 10k hours of bot-collected data @gen_intuition. It runs in real time at 20 fps, and posts highlighted a 5B-parameter model running an entire 2v2 match on a single NVIDIA B200, with no explicit physics or rendering engine @TheRundownAI. This was one of the clearest signals that video/world-model work is moving from toy demos toward interactive simulators.
- Speech remains highly competitive: AssemblyAI released Universal-3.5 Pro Realtime, a streaming STT model with 4.1% WER on AA-WER Streaming and contextual priming that can be updated mid-call without reconnecting @ArtificialAnlys. On the TTS side, Artificial Analysis said Speechify Simba 3.2 now leads its Speech Arena at 1233 Elo, ahead of Gemini 3.1 Flash TTS, Sonic 3.5, and Inworld Realtime TTS 1.5 Max, while also being the cheapest among top-ranked models @ArtificialAnlys.
- Document-context pipelines are becoming multimodal by default: LlamaIndex and LanceDB described a retrieval pipeline for messy PDFs that separates pages, chunks, and extracted assets into linked multimodal tables, reporting 82% any-page-hit@5 and 74% answer accuracy on a labeled ESG-report benchmark @lancedb, @llama_index. This pairs with Jerry Liu’s broader argument for a dedicated “document context layer” for agents @jerryjliu0.
Top tweets (by engagement)
- Anthropic’s global workspace paper dominated engagement, with the primary announcement on Claude’s internal workspace/J-space far above everything else @AnthropicAI.
- Tencent Hy3 was the biggest pure model-release story, especially among technical accounts discussing open-source competitiveness and deployment @teortaxesTex, @ShunyuYao12.
- MIRA’s playable world model was the standout multimodal/system demo @gen_intuition.
- Will Depue’s “Stargate for Data” thread was the most substantive strategy post, arguing that data collection—not compute alone—becomes the binding constraint and potential moat for frontier labs @willdepue.
- John Carmack’s memory-system thread drew significant technical interest by arguing inference hardware could exploit deterministic access patterns and much cheaper memory tiers than HBM for large-model serving @ID_AA_Carmack.
AI Reddit Recap
/r/LocalLlama + /r/localLLM Recap
1. Large Open-Weight MoE Model Releases
- longcat 2.0 (1.6T, ~48B active) weights are now open under MIT license (Activity: 638): LongCat 2.0 weights are now open under the MIT license via announcements from elie and ModelScope, with technical details in the LongCat 2.0 blog post. The model is a very large MoE system with 1.6T total parameters and roughly 48B active parameters per inference; commenters note the released weights occupy about 3.55 TB in BF16 and 2.05 TB in FP8. Commenters emphasized the practical deployment burden from the multi-terabyte weight size, and noted that Meituan—described as China’s Groupon/Uber Eats analogue—reportedly trained it on fully domestic Chinese chips, prompting discussion about the geopolitical/market significance.
Commenters highlighted the scale and deployment footprint of LongCat 2.0: 1.6T total parameters with approximately 48B active parameters, implying a sparse/MoE-style architecture. One user noted the released weights require about 3.55 TB in BF16 and 2.05 TB in FP8, which is important for anyone planning local storage or inference infrastructure.
- A technical point raised was that Meituan reportedly trained the model on 100% domestic Chinese chips, which commenters framed as significant for AI hardware supply-chain independence. This is especially notable given Meituan’s role as a major Chinese internet company comparable to a mix of Groupon and Uber Eats rather than a traditional AI lab.
- Several users focused on the permissive MIT license and planned benchmarking against frontier open models such as Qwen and DeepSeek. The combination of 1.6T total parameters, only ~48B active parameters, and open weights suggests the model may be practical to compare with other high-end MoE open models if inference tooling supports its architecture efficiently.
- New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0) (Activity: 604): Tencent released the non-preview Hy3 model collection on Hugging Face, described as a 295B-parameter MoE with 21B active parameters, now under Apache 2.0 rather than the prior restrictive community license. Commenters highlight a linked benchmark chart and claim the release shows “pretty impressive claimed gains over HY3-Preview”, potentially making it relevant for high-end local/home inference setups if real-world performance matches reported results. The main discussion is positive around the license change: commenters view moving from a geographically/restrictively limited license to Apache 2.0 as the most important improvement, especially alongside Tencent’s recent Apache-licensed translation models.
Commenters highlight that Tencent Hunyuan Hy3 is a large MoE-style release at 295B total parameters with 21B active, and one user notes the claimed benchmark gains over HY3-Preview appear substantial enough that, if they transfer to real workloads, it could be relevant for “high end home setups.” There is interest in practical inference availability, especially GGUF quantizations, which would determine whether local deployment is feasible.
- A technically important licensing change was noted: Tencent reportedly moved from a more restrictive “community” license that limited usage in regions such as South Korea, the UK, and the EU to Apache 2.0. Commenters view this as significant because it enables broader commercial and research reuse, consistent with some of Tencent’s recent Apache-licensed translation models.
- One commenter frames Hy3 as a potential alternative to Qwen and MiniMax, implying interest in whether its benchmark and real-world performance can compete with the current leading open-weight Chinese model families.
New model: GigaChat3.5-432B-A28B (with day-0 GGUF support!) (Activity: 439): Sberbank/ai-sage released GigaChat3.5-432B-A28B on Huggi
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み