高品質な人間データについて考える
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Lilian Weng
現代の深層学習モデル訓練において、高品質なデータは不可欠な燃料である。多くのタスク固有のラベル付きデータは、分類作業など人間による注釈付けから得られている。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
[Ian Kivlichan 氏(https://scholar.google.com/citations?user=FRBObOwAAAAJ&hl=en)に、100 年以上前の『Nature』誌に掲載された「Vox populi」という論文など、多くの有益な指摘と温かいフィードバックをいただき、心より感謝いたします。🙏]
高品質なデータは、現代の深層学習モデルのトレーニングにおける燃料です。タスク固有のラベル付きデータの多くは、LLM のアライメントトレーニングのための分類タスクや RLHF(人間による好みを基にした強化学習微調整)ラベリング(これは分類形式として構築可能)といった、人間の注釈によって提供されます。後述の多くの機械学習技術がデータ品質の向上に寄与しますが、根本的には高品質な人間のデータを収集するには、細部への注意と慎重な実行が不可欠です。コミュニティは高品質なデータの価値を理解していますが、どうやら「誰もがモデル作業をやりたがり、データ作業は避ける」という微妙な印象があるようです(Sambasivan et al. 2021)。
原文を表示
Special thank you to [Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. 🙏 ]
High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally human data collection involves attention to details and careful execution. The community knows the value of high quality data, but somehow we have this subtle impression that “Everyone wants to do the model work, not the data work” (Sambasivan et al. 2021).
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み