残差コンテキスト拡散言語モデル
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
Apple Machine Learning は、拡散大規模言語モデル(dLLMs)において、破棄されたトークンの計算を再利用する手法を示し、並列デコードの効率性を向上させる研究を発表した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
拡散大規模言語モデル(dLLMs)は、複数のトークンを並列にデコードできるため、純粋な自己回帰型言語モデルに対する有望な代替手段として登場しました。しかし、最先端のブロック分割型 dLLM は、「再マスク」メカニズムに依存しており、最も確信度の高いトークンだけをデコードして他を破棄するため、計算資源を実質的に浪費しています。これらの破棄されたトークンは、後のデコード反復において有用な文脈情報を保持しているため、その計算を再利用することは有益であることを示します。これを踏まえ、我々は残差文脈拡散(RCD)と呼ばれるモジュールを提案します。
原文を表示
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel. However, state-of-the-art block-wise dLLMs rely on a “remasking” mechanism that decodes only the most confident tokens and discards the rest, effectively wasting computation. We demonstrate that recycling computation from the discarded tokens is beneficial, as these tokens retain contextual information useful for subsequent decoding iterations. In light of this, we propose Residual Context Diffusion (RCD), a module that…
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み