手話モデルを用いた手話注釈の自己開始的生成手法
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Apple Machine Learning
研究者らは、高品質な手話データ不足という課題に対し、動画と英語を入力として候補注釈を自動生成する疑似注釈パイプラインを開発した。これにより、コストのかかる大規模注釈作業を軽減し、未利用のデータを活用可能にする。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI を駆使した手話解釈は、高品質な注釈付きデータの不足によって制限されています。ASL STEM Wiki や FLEURS-ASL といった新しいデータセットには専門的な通訳者や数百時間にわたるデータが含まれていますが、これらは部分的にしか注釈が付けられておらず、その規模での注釈作成にかかる莫大なコストが原因の一つとして、まだ十分に活用されていません。本研究では、手話動画と英語を入力とし、グロッス( glosses)、指文字、および手話分類器に対する可能性の高い注釈のランキング付きセット(時間間隔を含む)を出力する疑似注釈パイプラインを開発しました。このパイプラインは…
原文を表示
AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain professional interpreters and 100s of hours of data but remain only partially annotated and thus underutilized, in part due to the prohibitive costs of annotating at this scale. In this work, we develop a pseudo-annotation pipeline that takes signed video and English as input and outputs a ranked set of likely annotations, including time intervals, for glosses, fingerspelled words, and sign classifiers. Our pipeline uses sparse predictions from…
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み