残差コンテキスト拡散言語モデル(2 分読了)
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
最先端のブロック別拡散大規模言語モデルが、自信のあるトークンのみを復号し他を破棄する仕組みに対し、破棄されたトークンの情報を残差として次ステップに注入する新モジュール「Residual Context Diffusion」を開発した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
最先端のブロック形式拡散大規模言語モデル(dLLMs)は、最も確信度の高いトークンのみを復号化し、それ以外のトークンを破棄する再マスク機構に依存しています。破棄されたトークンから計算を再利用することは有益です。なぜなら、これらのトークンは後続の復号化反復において有用な文脈情報を保持しているからです。残差文脈拡散(Residual Context Diffusion)は、このように破棄されたトークン表現を文脈上の残差に変換し、次のノイズ除去ステップに再注入するモジュールです。これは、広範なベンチマークにおいて、最小限の追加計算オーバーヘッドで最先端の dLLMs の精度を一貫して向上させます。
原文を表示
State-of-the-art block-wise Diffusion Large Language Models (dLLMs) rely on a remasking mechanism that decodes only the most confident tokens and discards the rest. Recycling computation from the discarded tokens is beneficial, as these tokens retain contextual information useful for subsequent decoding iterations. Residual Context Diffusion is a module that converts these discarded token representations into contextual residuals and injects them back for the next denoising step. It consistently improves frontier dLLMs in terms of accuracy with minimal extra computation overhead across a wide range of benchmarks.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み