現在の言語モデル学習はインターネットの大部分を活用できていない
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Decoder
アップル、スタンフォード大学、ワシントン大学の研究者らが、HTML抽出ツールの選択によって言語モデルの学習データが大きく異なり、ウェブコンテンツの大部分が活用されていないことを発見した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

大規模言語モデルはウェブデータから学習しますが、どのページが実際にトレーニングセットに含まれるかは、一見些細な選択であるHTMLエクストラクター(HTML抽出ツール)に大きく依存しています。Apple、スタンフォード大学、ワシントン大学の研究者らは、3つの一般的な抽出ツールが同じウェブページから驚くほど異なるコンテンツを抽出することを発見しました。
この記事「Current language model training leaves large parts of the internet on the table」は、The Decoderで最初に公開されました。
原文を表示

Large language models learn from web data, but which pages actually make it into training sets depends heavily on a seemingly mundane choice: the HTML extractor. Researchers at Apple, Stanford, and the University of Washington found that three common extraction tools pull surprisingly different content from the same web pages.
The article Current language model training leaves large parts of the internet on the table appeared first on The Decoder.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み