Google DeepMind、聴覚障害者向け手話 AI モデル「SL2T」を発表
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Google DeepMind
Google DeepMind は、聴覚障害者向けに多言語手話テキスト変換モデル SL2T を開発し、Pixel 11 の Gboard や Live Transcribe で実装したと発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 00:47
AI深層分析
キーポイント
SL2T モデルの発表と実用化
Google DeepMind は手話テキスト変換(SL2T)モデルを公開し、初めて消費者向け製品に組み込むことで実験室から実社会へ技術をもたらした。
Pixel 11 における新機能
アメリカ手話(ASL)から英語への翻訳機能を Gboard と Live Transcribe に搭載し、Web 検索やメッセージ作成を指文字で行えるようにする。
聴覚障害コミュニティへの影響
手話は聴覚障害者の文化的アイデンティティの核心であり、この技術は両コミュニティ間のコミュニケーションギャップを埋める新たな可能性を開く。
手話翻訳の核心的な課題
手話は音声言語とは異なり独立した自然言語であり、単なる逐語変換ではなく文法と語彙を持つ真の意味での機械翻訳を必要とする。また、意味は手、腕、体幹、顔の同時的な動きで伝達されるため、高精度なコンピュータビジョン技術が不可欠である。
SL2T の学習アプローチとデータ規模
モデルは50種類以上の手話言語にわたる10万時間超のデータで訓練され、多様な言語や方言を同時に学ぶことで単一言語モデルを上回る性能を発揮する。
重要な引用
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
We are bringing sign language AI out of the lab and into consumer products for the first time
According to our testers, signing in ASL is faster, more natural, and more delightful than typing in English.
"Sign languages aren't simply 'English on the hands.'"
編集コメントを表示
編集コメント
聴覚障害者の言語的・文化的権利を尊重し、技術的な障壁を取り除くための重要な一歩である。Google は手話の多様性を認識しつつ、まずは主要な手話から実装を進める戦略を採用している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
2026 年 8 月 12 日
Google DeepMind 手話チーム
手話からテキストへ変換する「SL2T」モデルをご紹介します。これは、聴覚障害者や難聴者のユーザー向けに新機能を搭載した画期的なモデルです。
過去数十年で AI の音声処理能力は急速に進化し、聴覚のある人にとっては自然な自動翻訳、音声入力、対話型インターフェースが実現しました。しかし、この技術的ブレイクスルーはまだ、世界中の 200 を超える手話や、それらを使用する推計 7,000 万人の聴覚障害者・難聴者の世界には届いていませんでした。
本日、私たちは品質と汎用性において画期的な成果を収めた、大規模多言語対応の手話からテキストへ変換(SL2T)翻訳モデルを発表します。これにより、手話 AI が研究の枠を超え、初めて一般消費者向け製品に実装されることになります。具体的には、Gboard と Live Transcribe で「手話入力」機能が利用可能になります。対象は Pixel 11 のアメリカン・サイン・ランゲージ(ASL)から英語への変換です。対応端末や言語の拡充も間もなく予定されています。
聴こえる人が音声入力を使ってタイピングの代わりに話すように、この機能を使えば、聞こえない人も普段文字を入力する場所でどこでも手話で操作できます。ウェブ検索を手話で行ったり、メッセージや文書の作成を手話で始めたり、Gemini に質問したりタスクを実行させたりすることも可能です。また「ライブ・トランスクリプト」では、会話の手応えをタイピングでやり取りする代わりに、手話で返答できます。テストユーザーによると、アメリカ手話(ASL)での入力は、英語のタイピングよりも速く、自然で、より楽しいものです。
SL2T を基盤とした「Sign-to-text」機能により、聞こえない人は普段文字を入力する場所でどこでも手話で操作できるようになります。
なぜ手話が重要なのか
手話は世界中の聴覚障害者コミュニティにおける主要な言語であり、聴覚障害者の文化的アイデンティティの中核を成すものです。手話や発話、読書、筆記の習熟度には個人差が大きく存在するため、あらゆるコミュニケーション手段へのアクセスを支援することが重要です。聴覚障害者は、聴こえる人が音声処理技術から恩恵を受けるのと同様に、手話処理技術からも利益を得られます。さらにこの技術は、聴覚者と聴覚障害者コミュニティ間のコミュニケーションギャップを埋める新たな可能性も開きます。
こうした社会的インパクトの可能性にも関わらず、手話 AI の進展は緩やかです。その理由の一つは、手話向け AI を構築する際に複雑な課題に直面すること、もう一つの理由は、手話そのものの仕組みについて広く誤解が存在していることです。
音声言語の文字起こしと比較すると、手話翻訳には2つの核心的な課題があります。まず、音声の文字起こしは同じ言語内で音声をテキストへ逐次的に変換する作業ですが、手話は独自の文法と語彙を持つ独立した自然言語です。そのため、単なる「記号→単語」の変換プロセスではなく、真の意味での機械翻訳が求められます。
2つ目の課題は、モデルが物理的な動きを「見て」理解できる必要がある点です。手話では、手や腕、体幹、頭部、顔の同時的な動きによって意味が伝達されます。これらを高いフレームレートで正確に追跡するのは、計算コストが高く難しいコンピュータビジョンタスクです。
こうした背景を踏まえると、手話グローブのような初期の手話技術がなぜ根本的に限界を持っていたかが理解できます。手話は単なる「手の英語」ではないからです。微細な全身の動きに対する複雑な視覚的知覚と、本格的な言語翻訳能力が必要です。SL2T は、この両方を提供するために設計されています。
SL2T は、手話の入力を signer の身体上のポイントとして捉え、ストリーミング形式でテキスト出力に変換します。FLEURS-ASL ベンチマークからの例。
SL2T の仕組み
SL2T は、ユーザー中心かつ文化的背景を考慮したアプローチと、大規模なデータスケーリングを組み合わせて構築されました。このモデルは、50 以上の手話言語にわたる 10 万時間以上ものデータを学習対象としており、そのうち約四分の一が米国手話(ASL)のデータです。多様な言語や方言、習熟度を同時に学習させることで、モデルは共通する根本的な構造を把握し、単一言語モデルを上回る性能を発揮することが実験で確認されています。
ユーザーのプライバシー保護のため、SL2T は生カメラ映像そのものではなく、ポーズのランドマーク(特徴点)の位置シーケンスとして手話を扱います。オンデバイス上で動作する MediaPipe Holistic が、話者の身体上のポイント位置を追跡し、サーバーに送られるのはこれらの幾何座標のみです。これにより、元の映像は即座に破棄され、プライバシーが守られます。
SL2T はこの座標シーケンスを直接テキストに変換します。従来の手話翻訳研究で広く用いられていた中間的な注釈である「グロッス(glosses)」を経由しないのが特徴です。グロスでは、非手指マーカーや空間構文など、手話の豊かで非線形的な側面を捉えきれないという課題がありました。ランドマークから直接翻訳を行うことで、人工的な語彙制限を取り払い、データ量に応じて翻訳品質が直線的に向上する仕組みを実現しています。
FLEURS-ASL (sd-test) などの主要なベンチマークにおいて、SL2T は現在までに存在する中で最も能力の高い手話翻訳モデルです。このモデルは ASL から英語への翻訳品質を評価する基準で、70 BLEURT という驚異的なゼロショットスコアを達成しており、これまでに報告されたどのスコアよりも大幅に上回っています。
しかし、学術的なベンチマークの最適化だけでは、実世界でのアプリケーションにおける使いやすさを保証することはできません。そのため、私たちはストリーミング遅延の最小化、非手話入力に対するハルシネーション(幻覚)の防止、左利きの手話話者(約 10%)への公平性の確保、そして片手でスマートフォンを保持しながら行う片手での手話における性能向上といった、実用的な課題に取り組んできました。
| 元の英語 | SL2T の ASL → 英語出力 |
|---|---|
| クック諸島には都市がありませんが、15 の異なる島で構成されています。主な島はラロトンガとアイトウタキです。 | クック諸島には都市がなく、15 の島から成り立っています。主要な 2 つの島はラロトンガとアイトウタキです。 |
| 試合は午前 10:00 に好天の下で開始され、午前中の小雨がすぐに上がった以外は、7 ラグビーに完璧な一日でした。 | 試合は午前 10 時に好天の下で開始されます。朝に軽い雨が降りますがすぐに上がります。7v7 ラグビーには完璧な一日です。 |
| 米国やカナダなどの一部の連邦国家では、所得税は連邦レベルと地方レベルの両方で課税されるため、税率や控除枠は地域によって異なります。 | 米国やカナダなどの一部の連邦国家では、所得税は連邦および地方レベルの両方で徴収されます。つまり、税率や控除枠は地域によって異なります。 |
| この全身に羽毛が生え、恒温性の猛禽類は、ヴェロキラプトルのような爪を持ち、2 本の足で直立して歩いていたと考えられています。 | この生物は恒温性で灰色を食べ、羽毛に覆われています。ヴェロキラプトルのように 2 本の足で歩くと考えられています。 |
| いつか、あなたのひ孫たちが異星の世界の頂上に立ち、古代の祖先について疑問に思うようになるかもしれません? | いつか、あなたの子孫たちが異星の世界に立ち、祖先について振り返るようになるかもしれません。 |
FLEURS-ASL ベンチマークからの例。SL2T は複雑な手話を流暢な英語に正確に変換しますが、稀なサインや高速の指文字("prey" が "grey" となるなど)、受動態、分類詞による描写(「爪」が省略されるケース)、文脈のない時制表現("kicked off" が "start" となるなど)では、依然として偶発的な誤りが発生します。
コミュニティと共に構築する
私たちは、デフコミュニティのために作るのではなく、彼らと共に作ること信じています。このプロジェクトのあらゆる段階——Google のデフ社員である Sam Sepah 氏による概念設計から、デフパートナーとのデータ収集、デフユーザーによる評価、そして技術の影響をデフ専門家と評価に至るまで——すべてにデフの視点が反映されています。
責任ある実世界での導入を導くため、「AI 手話諮問委員会(AISLAC)」を設立し、世界中の多くのデフ団体や専門家を結集させました。この参加型ガバナンスモデルを通じて、技術の影響を最も受けるコミュニティが、開発の優先順位に直接影響を与えています。Gboard と Live Transcribe における SL2T 1.0 のリリースに合わせて共同で作成した インパクトレポート では、技術の能力と現在の限界を透明性を持って詳述しています。これは主要な手話言語リリースにおいて継続していく方針です。
今後の展望
SL2T は、学界と産業界で数十年にわたって積み重ねられてきた基盤研究の上に成り立っていますが、米国手話(ASL)の入力をユーザーのスマートフォンへ届けることは、その始まりに過ぎません。Google のミッションは「世界の情報を整理し、誰もがアクセスできて有用なものにすること」です。真の意味での普遍的なアクセシビリティを実現するには、音声や文字言語と同等のレベルを達成する必要があります。当チームは今後、この技術を他の手話言語への対応、手話生成機能、そして最先端の AI 能力へと拡張していく所存です。責任ある形で進捗を発表し、デジタル空間全体で手話によるアクセスが標準となるよう努めてまいります。
SL2T はまず Pixel 11 の Gboard と Live Transcribe で体験いただけます。今後対応端末も順次追加されますが、いずれも追加費用はかかりません。
謝辞
本プロジェクトは Google DeepMind と Android チームの共同作業によって実現されました。SL2T モデルの開発に携わったコアメンバーは以下の通りです:Garrett Tanzer, Benoit Brard, Elizabeth Clark, Tim Dozat, Sebastian Ebert, Dan Garrette, Manfred Georg, Vicky Holgate, Shankar Kumar, Mohammad Saboorian, Miloš Stanojević, Megh Umekar, John Wieting, Andy Zhang, Chris Dyer。
Gboard および Live Transcribe へのモデル統合を担当した Android チームは以下のメンバーです:Ausmus Chang, Sai Aditya Chitturu, Dayle Chiu, Anna Chou, Ajay Dudani, Angana Ghosh, Alex Huang, Joanne Kim, Ed Lee, Thomas Lin, James Su, Yanchao Su, Sharlene Yuan。
本稿の作成にあたり、Anelia Angelova 氏、Abhishek Bapna 氏、Sara Basson 氏、Glenn Cameron 氏、Scott Crowell 氏、Trevor Cohn 氏、Noah Fiedel 氏、Zoubin Ghahramani 氏、Raia Hadsell 氏、Tom Hudson 氏、Alexander Hauerslev Jensen 氏、Kazuya Kawakami 氏、Peike Li 氏、Liam McCafferty 氏、Caroline Pantofaru 氏、Abhinav Parashar 氏、Christopher Patnoe 氏、Laura Rimell 氏、Sam Sepah 氏、Thad Starner 氏、Dave Uthus 氏、Biao Zhang 氏の皆様に、多大なるご支援を賜りましたことを心より感謝申し上げます。
また、モデルの初期段階におけるテストにご協力いただいた皆様にも、厚く御礼申し上げます。
原文を表示
August 12, 2026 Models
Google DeepMind Sign Language Team
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
AI's ability to process spoken languages has advanced rapidly over recent decades, enabling automatic translation, dictation, and conversational interfaces that feel effortless to hearing users. Yet this technological revolution has not reached the world’s more than 200 sign languages — and the estimated 70 million Deaf and hard of hearing people who use them.
Today, we’re introducing a massively multilingual sign-language-to-text (SL2T) translation model that marks a breakthrough in quality and generality. With it, we are bringing sign language AI out of the lab and into consumer products for the first time: SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English. More devices are coming soon, and additional languages will follow.
Similarly to how hearing users can use dictation to speak instead of typing, this feature enables Deaf users to sign to their phone anywhere they’d normally type. You can sign to search the web, draft messages or documents, and ask Gemini to solve queries or execute tasks. In Live Transcribe, you can sign responses in conversations instead of having to type back and forth. According to our testers, signing in ASL is faster, more natural, and more delightful than typing in English.
Sign-to-text, powered by SL2T, enables users to sign to their phone anywhere they'd normally type.
Why sign languages matter
Sign languages are the primary languages of Deaf communities around the world and the cornerstone of Deaf cultural identity. There is great diversity among deaf people in terms of their level of proficiency in signing, speaking, reading, and writing, so it is important to support access in all modalities. Deaf people can benefit from sign language processing in the same way that hearing people benefit from spoken language processing, plus the technology opens new possibilities for bridging the communication gap between Deaf and hearing communities. Despite this opportunity for positive social impact, progress in sign language AI has been slow — both because building AI for sign languages presents complex challenges and because widespread misconceptions exist about how the languages themselves work.
Compared to spoken language transcription, sign language translation presents two core challenges. First, transcribing speech is a matter of performing a sequential mapping from sound to text in the same language, whereas sign languages are independent, natural languages with their own distinct grammars and lexicons. As a result, they require true machine translation rather than a sequential process of sign-to-word transformations. Second, the model must learn to “see” and understand physical movement. Sign languages convey meaning through simultaneous movements of the hands, arms, torso, head, and face. Accurately tracking these at high frame rates is a difficult and computationally demanding computer vision task.
Given this background, it is easy to understand why some early attempts at sign language technology, like sign language gloves, were fundamentally limited: sign languages aren't simply “English on the hands.” They require complex visual perception of fine-grained whole-body movements and full-fledged language translation. SL2T is designed to deliver both.
SL2T sees sign language inputs as points on the signer's body and translates them into streaming text outputs. Example from the FLEURS-ASL benchmark.
How SL2T works
We built SL2T by combining a user-centric, culturally informed approach with massive data scaling. The model is trained on over 100,000 hours of data across more than 50 sign languages — with roughly a quarter of the data in ASL. Training jointly on diverse languages, dialects, and proficiency levels causes the model to learn shared underlying structures, outperforming single-language models in our experiments.
To protect user privacy, SL2T sees sign language as a sequence of pose landmark locations rather than a raw camera feed. An on-device model (MediaPipe Holistic) tracks the location of points on the signer, and only these geometric coordinates are sent to the server for translation, allowing the original video to be discarded immediately.
SL2T translates this coordinate sequence directly into text, bypassing intermediate annotations known as “glosses” that are widely used in prior work on sign language translation. Glosses fail to capture rich, non-linear aspects of sign languages such as non-manual markers and spatial constructions. Translating directly from landmarks removes artificial vocabulary limits and allows translation quality to scale directly with data.
SL2T is the most capable sign language translation model to date according to key benchmarks like FLEURS-ASL (sd-test), which assesses ASL to English translation quality. SL2T achieves a remarkable zero-shot score of 70 BLEURT, which is significantly higher than any previously reported score. But optimizing academic benchmarks alone doesn’t guarantee usability in real-world applications, so we worked hard on practical issues like minimizing streaming latency, preventing hallucination on non-signing inputs, ensuring fairness for the 10% of signers who are left-handed, and improving performance for one-handed signing, which is used while holding a smartphone in the other hand.
| Original English | SL2T's ASL → English output |
|---|---|
| The Cook Islands do not have any cities but are composed of 15 different islands. The main ones are Rarotonga and Aitutaki. | The Cook Islands have no cities and consist of 15 islands. The two main islands are Rarotonga and Aitutaki. |
| The games kicked off at 10:00am with great weather and apart from mid morning drizzle which quickly cleared up, it was a perfect day for 7's rugby. | Games start at 10 a.m. in great weather. There is a light rain in the morning that clears up. It's a perfect day for 7v7 rugby. |
| In some federal countries, such as the United States and Canada, income tax is levied both at the federal level and at the local level, so the rates and brackets can vary from region to region. | In some federal countries, like the US and Canada, income tax is collected at both the federal and local levels. This means that the rates and brackets vary depending on your region. |
| This fully feathered, warm blooded bird of prey was believed to have walked upright on two legs with claws like the Velociraptor. | This creature is warm-blooded, eats grey, and is covered in feathers. It is believed that it walks on two legs like a velociraptor. |
| Maybe one day, your great grandchildren will be standing atop an alien world wondering about their ancient ancestors? | Maybe one day your great-grandchildren will stand on an alien world and reflect on their ancestors. |
*Examples from the FLEURS-ASL benchmark. SL2T accurately translates complex ASL into fluent English. Occasional errors remain in rare signs, rapid fingerspelling ("prey" → "grey"), passive constructions, classifier depictions (dropping "claws"), and tense without context ("kicked off" → "start").*
Building with the community
We believe in building *with* the Deaf community, not just *for* it. Deaf perspectives have shaped every stage of this project — from conceptualization by Sam Sepah, a Deaf Googler, to data collection with Deaf partners, evaluation in Deaf user studies, and impact assessment of the technology with Deaf experts.
To guide responsible real-world deployment, we established the AI Sign Language Advisory Committee (AISLAC), bringing together many global Deaf organizations and subject-matter experts. Through this participatory governance model, the communities most impacted by our technology directly influence our development priorities. We co-authored a joint impact report for the release of SL2T 1.0 in Gboard and Live Transcribe, transparently detailing the technology's capabilities and current limitations — a collaborative approach we plan to continue for all major sign language releases.
Looking ahead
SL2T builds upon decades of foundational research across academia and industry, but bringing ASL input to users’ phones is only the beginning. Google’s mission is to organize the world's information and make it universally accessible and useful. Achieving universal accessibility means reaching full parity with spoken and written languages. Our team is working to expand this technology into additional sign languages, sign language generation, and frontier AI capabilities. We look forward to sharing our progress responsibly in order to make access through sign languages standard across the digital landscape.
You can experience SL2T in Gboard and Live Transcribe first on Pixel 11, with more devices coming soon — all at no additional cost.
Acknowledgements
This work was done jointly by teams from Google DeepMind and Android. The core team who developed the SL2T model is: Garrett Tanzer, Benoit Brard, Elizabeth Clark, Tim Dozat, Sebastian Ebert, Dan Garrette, Manfred Georg, Vicky Holgate, Shankar Kumar, Mohammad Saboorian, Miloš Stanojević, Megh Umekar, John Wieting, Andy Zhang, and Chris Dyer.
The Android team who integrated the model into Gboard and Live Transcribe is: Ausmus Chang, Sai Aditya Chitturu, Dayle Chiu, Anna Chou, Ajay Dudani, Angana Ghosh, Alex Huang, Joanne Kim, Ed Lee, Thomas Lin, James Su, Yanchao Su, and Sharlene Yuan.
We are grateful for additional support from Anelia Angelova, Abhishek Bapna, Sara Basson, Glenn Cameron, Scott Crowell, Trevor Cohn, Noah Fiedel, Zoubin Ghahramani, Raia Hadsell, Tom Hudson, Alexander Hauerslev Jensen, Kazuya Kawakami, Peike Li, Liam McCafferty, Caroline Pantofaru, Abhinav Parashar, Christopher Patnoe, Laura Rimell, Sam Sepah, Thad Starner, Dave Uthus, and Biao Zhang.
Many thanks also go to those who participated in early stage testing of our models.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み