物理的 AI の次なる鍵は脳波か?技術の可能性を探る
Encord は Zander Labs と提携し、人間の脳波を計測して作業中の意図やエラーを検知するデータセットの構築を試みており、物理的 AI の学習に必要な実世界データの不足というボトルネック解消を目指す。
AI深層分析を開く2026年7月27日 13:03
キーポイント
脳波を活用したデータ収集の試み
Encord は Zander Labs が開発したヘッドセットを用い、作業者がブロックを組む際の脳波を計測し、エラーや意図といったメンタル状態をタグ付けしたデータを生成している。
物理的 AI のデータ不足というボトルネック
Encord の責任者によると、ロボット工学の次なる制約はモデルアーキテクチャではなく、実世界の物理トレーニングデータの絶対的な不足にあると指摘されている。
生成 AI による解決への限界認識
自動運転車や動画からの学習といった既存のアプローチでは、実世界データの忠実度(fidelity)に欠けるため、物理的タスクにおける完全な学習には不十分であると分析されている。
パイロット試験と評価プロセス
同社はまず脳波タグ付きデータセットを構築し、顧客のロボットモデルで動作テストを行った上で、性能向上が確認できた場合にのみ本格的な展開を検討する方針を示している。
物理的AIの突破には膨大なデータセットが必要
Velmurugan氏によると、現在の壁を突破するにはYouTubeの動画コレクションの約5倍規模のデータセットが必要とされる。この巨大なスケールが、データ生成自体が研究課題からビジネスへと変化した理由となっている。
重要な引用
The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California.
Encord says the goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up.
The data simply does not exist.
"Velmurugan says it will take a data set something like five times the size of YouTube's video corpus to break through"
編集コメントを表示
編集コメント
ロボット学習におけるデータ不足という課題に対し、人間の脳活動という新たな次元の情報を活用する試みは極めて革新的である。ただし、現時点ではパイロット試験段階であり、実用化への道筋はまだ不透明な部分が残っている。
物理 AI の最前線は、カリフォルニア州サンレアンドロの倉庫で行われる「ジャングラ」ゲームのようなものです。
その倉庫には、AI モデルの訓練に用いるデータツールを開発する企業 Encord が進出しています。同社ではロボットトレーナーを「パイロット」と呼んでいますが、アンソニー・セージャ氏もその一人です。彼はカメラ付きヘッドセットを着用し、崩れかけたタワーから木製のブロックを慎重に取り外しています。これはロボット訓練データの収集において決して珍しいことではありませんが、このヘッドセットには脳波を計測するセンサーも搭載されています。彼がブロックの塔を組み立てる過程で、その脳波が記録されるのです。
Encord は、ヒューマノイドや倉庫用ロボットの今後の最大の制約要因はモデルアーキテクチャではなく、現実世界での物理訓練データの圧倒的な不足にあると見なす、少数ながら増加しつつあるスタートアップの一つです。同社は既存のデータ管理を支援するだけでなく、自社で欠けているデータを製造するというビジネスモデルを構築しています。
セージャ氏が装着している脳波計測ヘッドセットは、ドイツの神経科学スタートアップである Zander Labs が開発したものです。同社は、エラーや意図、驚きといった精神状態を推定するために脳活動を測定することで、より有用なデータセットを作成できると考えています。Encord と Zander の共同作業は現在、試験段階にあります。Encord によると、まずは脳波タグ付きの初期データセットを構築し、顧客が持つロボットモデルに適用して性能向上を実証した上で、本格展開の可否を判断する計画です。
Zander の神経科学者で、この研究を監督しているルカス・ゲルケは、特定のタスクのどの時点でどれだけの脳活動が必要かというデータが、モデル開発者が「いつ最高性能のモデルを投入すべきか」を見極めるための手がかりになると語っています。
ロボット学習担当責任者のヴァイネース・ベルムラガンによれば、これはロボティクスにおけるデータボトルネックを解消するための取り組みにおいて、「最前線(ブリーディング・エッジ)」に位置するものです。OpenAI のロボットラボや倉庫自動化企業 Berkshire Grey で経験を積んだベテランである彼は、Encord 入社後は同社内部のデータ作成チーム構築を担当しています。
Encord は、機械学習ビジョンアプリケーションを開発する企業がデータを注釈付けし、モデルを評価できるよう支援するために設立されました。ベルムラガンによると、同社は多くの一流ロボット企業と取引していますが、名前は明かせないとのことです。これらの顧客がロボットの操作タスクにエンドツーエンドの学習を適用し始めた際、経営陣は単にデータを管理するだけでなく、自らトレーニングデータを作成しなければならないことに気づきました。「既存のデータでは足りないのです」とベルムラガンは言います。
生成 AI がチャットボットで成し遂げたことをロボットにも実現できるという賭けは、相変わらず同じ壁にぶつかり続けています。自動運転車メーカーは自前で実世界データを収集していますが、それをスケールさせるのは困難です。動画からの学習も可能ですが、実世界のデータには及ばない忠実度しか得られません。Velmurugan 氏によれば、この壁を突破するには YouTube の動画コレクションの約 5 倍規模のデータセットが必要だというのです。その規模感こそが、データ生成自体が研究課題からビジネスへと進化している理由を説明しています。
エゴセントリックなデータのニーズに応える
ロボット用「脳」を開発する企業は、現在主に 2 つのデータソースに注目しています。1 つ目は作業員が装着したカメラで撮影した「エゴセントリック(自己視点)」動画です。必要に応じて複数のカメラアングルや他の計測データを追加することもあります。もう 1 つは遠隔操作されたロボットから得られるデータです。
Encord は両方のアプローチを採用しています。同社は世界各地の工場からエゴセントリックデータを収集し、サンレアンドロにある自社工場では脳波のような新しいモダリティの実験や、特定のスキルに特化した微調整用データセットの収集を行っています。
TechCrunch が訪問した際、パイロットプロジェクトでは「リーダー・フォロワー」方式が採用されていました。これは 2 本のロボットアームをペアにした仕組みで、1 本は人間オペレーターが直接操作し、もう 1 本はその動きを模倣します。これによって、コーヒーポットからマグカップへ注ぐ(液体の揺れが激しい)や、ポーカーチップを積み重ねるといったタスクに関するデータを生成しています。「ヒューマノイド型ロボットを開発する企業は皆、こうしたデータセットを求めてきます」と Velmurugan 氏は語ります。
ストレージラックには、人工花の鉢植えや本、プラスチック製の野菜、猫用トイレとスコップ、ワイヤーの袋や束などが積み上げられていました。これらは家庭内タスク用のマニピュレータを訓練するための素材です。
この作業場の一つで、もう一人のパイロットであるソフィア・インファンテは、ロボットアームを使ってサーバー背面のイーサネットケーブルを抜き差ししていました。データセンターのオペレーターが望むような自動化ですが、ロボットが必要な精度で操作できればの話です。私は実際にコントローラーを握って操作を試みましたが、なぜまだ実現できていないのか理解できました。ピンセットのような形状は人間の指ほど器用ではなく、人間が当然のように持っている腕の自由度も欠いているのです。
Encord が開発中の新しいデータモダリティでは、前腕に装着したセンサー群を使って筋肉内の電気信号を検出します。通常、物体を操作する人間の手の動画撮影では手全体が映りきらないものですが、ヴェルムラガン氏はこれらのアームセンサーから、任意の瞬間における手の位置を 3D で描画する仕組みを構築したいと考えています。これにより、モデルに対するより堅牢な理解をもたらすことが期待されます。
Encord のデータセットには、「右手でボルトを締める」のように各動画の内容を物理的に記述した注釈が付けられています。これは LLM ベースのモデルが何が起こっているかを理解するのを助けるためです。ヴェルムラガン氏によると、こうした詳細な注釈は、特定のタスク訓練用の「質の低い自己データ(junky ego data)」に比べて 100 倍もの価値があり、その生産コストも 20 倍程度で済むため、机上の計算では非常に有利なトレードオフになるとのことです。
しかし「20 倍」という数字は、決して無視できない金額です。ここが最大の課題です。LLM の開発者が Stack Overflow やウェブ上のテキストを収集してモデルを構築した際、そのコストは最前線の研究機関にとってほぼゼロに等しかったからです。一方、物理的な AI の学習データを生成するには莫大な費用がかかります。これが「物理 AI を LLM と単純比較する」ことの限界でもあります。この種のデータは単に集めるだけでは不十分で、製造する必要があります。そのため、モデル構築における経済構造そのものが変わってしまうのです。
ベルムルガン氏は、業界全体で進歩が起きていると語ります。Encord が業界各社のプログラムを可視化できる立場にあるため、スタートアップから最前線の研究機関まで、何が有効で何がそうでないかを検証し、物理 AI モデルの改善に取り組んでいる様子を直接見ることができます。このように複数のロボット企業にまたがる視点を持つこと自体が、Encord の強みです。特定の顧客が気づくよりも早く、業界全体でどのデータ技術が注目されつつあるかを察知できるのです。
これにより、Encord の施設で行われている十数件のパイロットプロジェクトは常に活気に満ちています。インファント氏とセハ氏はともに、ニューラルネットワークの基礎を構築する新たな労働力の一部です。彼らは Encord 入社前に、AI データ注釈を手掛けるもう一つの企業である Scale で働いていました。
セハ氏は以前、廃棄物管理会社で働き、そこで技術への関心からロボットごみ選別機の稼働維持を担当していました。今では積み上げられたブロックが崩れ落ちるような難題に直面する中で、ロボットの学習タスクを解決する挑戦を楽しんでいます。「毎日新しい発見があります!」と彼は言います。
当記事内のリンクを通じてご購入いただいた場合、小規模な手数料をいただいております。ただし、これは当社の編集の独立性には一切影響しません。
原文を表示
The frontier of physical AI is a Jenga game in a warehouse in San Leandro, California.
That warehouse is occupied by Encord, a company that builds data tooling used to train AI models. Andrew Ceja is a pilot—the company’s term for its robotic trainers—and he’s carefully pulling wooden blocks from a tottering tower while wearing a headset with a camera that tracks what he sees. That alone is fairly common for collecting robot training data, but this headset includes sensors that measure his brain waves as he carefully disassembles the block tower.
Encord is one of a small but growing number of startups betting the next real constraint on humanoid and warehouse robotics won’t be model architecture but instead the sheer scarcity of real-world physical training data. Rather than just helping robotics companies manage the data they have, Encord is building a business around manufacturing the data they don’t.
The brain wave headset Ceja is wearing was built by Zander Labs, a German neuroscience startup that’s betting measuring brain activity — to deduce mental states like error, intent and surprise — can create a more useful data set to train models. Encord’s work with Zander is currently a trial run; Encord says the goal is to build an initial brain wave-tagged data set, run it through customer robotics models, and evaluate whether it actually improves performance before deciding whether to scale it up.
Lucas Gehrke, a Zander neuroscientist supervising the work, says that the amount of brain activity used at any point during a given task offers clues for model builders trying to figure out when they need to deploy their highest-effort models.
This is the “bleeding edge” of the effort to solve the robotics data bottleneck, according to Vineeth Velmurugan, Encord’s head of robot learning. A veteran of OpenAI’s robot lab and Berkshire Grey, the warehouse automation firm, Velmurugan joined Encord to build the company’s internal data-creation team.
Encord was founded to help companies building machine-vision applications annotate data and evaluate models. As their customers—Velmurugan says they work with many leading robotics firms but that he’s not authorized to name them—began to apply end-to-end learning to robotic manipulation tasks, executives realized they would have to produce training data themselves, rather than simply manage it. “The data simply does not exist,” Velmurugan said.
The bet that generative AI can do for robots what it’s done for chatbots keeps running into this same wall. Self-driving car companies collect physical-world data themselves, but that’s hard to scale. Training from video can work, but it lacks the fidelity of real world data. Velmurugan says it will take a data set something like five times the size of YouTube’s video corpus to break through—a scale that helps explain why data-generation itself has become a business and not just a research problem.
Feed your egocentric data needs
Companies building robot brains are now turning to two main sources: “egocentric” video collected by workers wearing cameras, often augmented with additional camera angles and other metrics, and data from robots operated remotely. Encord does both, drawing egocentric data from several factories around the globe, and using its San Leandro facility to experiment with new modalities, like brain waves, or collect data sets around specific skills for fine-tuning.
When TechCrunch visited, pilots were using leader-follower rigs — paired robotic arms, one controlled directly by a human operator and one that mimics its movements —to create data about tasks like pouring coffee from a pot into mugs (very sloshy) and stacking poker chips. “Every humanoid company has asked us for these pieces,” Velmurugan says.
Storage racks held cartons of fake flowers in vases, books, plastic vegetables, kitty litter trays and scoops, bags and bundles of wires, the stock in trade for training manipulators for household tasks.
At one of these stations, another pilot, Sofia Infante, maneuvers robotic arms to plug and unplug ethernet cables from the back of a server—the kind of work data center operators would love to be automated, if only robots could manipulate them with the required precision. Taking a spin behind the controls, I was able to see why that’s still out of reach: Pincers are far less dextrous than human fingers and lack the degrees of freedom we take for granted in our arms.
Another new data modality that Encord is developing uses a set of sensors strapped to the forearm to detect electrical signals in muscles. Video taken of human hands manipulating objects typically doesn’t capture the entire hand, but Velmurugan hopes to build a 3D depiction of where the hand is at any time based on the arm sensors, creating a more robust understanding for models.
Encord’s data sets are annotated with physical descriptions of what each video contains—”right hand tightens bolt”—to aid LLM-based models in understanding what is happening. Velmurugan estimates this kind of dense annotation is worth 100 times as much as “junky ego data” for training specific tasks, and it only costs 20 times more to produce, which is a good trade, on paper.
But “20 times more” is still real money, and that’s the catch: scraping text off the internet, the way LLM makers built their models by pulling from Stack Overflow and the rest of the web, cost frontier labs next to nothing. Generating physical training data does not, and that’s the limit of the physical-AI-as-LLM comparison. This kind of data has to be manufactured, not just collected, and that changes the economics of building these models.
Velmurugan says that progress is being made—with Encord’s visibility into programs across the industry, he’s able to see start-ups and frontier labs alike figure out what works and what doesn’t to improve physical AI models. That vantage point—sitting between many robotics companies at once—is also part of Encord’s pitch. It can spot which data techniques are gaining traction industry-wide before any single customer can.
That will keep the dozen or so pilots at Encord’s facility busy. Both Infante and Ceja are part of a burgeoning workforce developing the building blocks for neural networks; they previously worked at Scale, another AI data annotation firm, before joining Encord.
Ceja had worked at a waste management company where his interest in technology found him in charge of keeping a robotic trash sorter in good working order. Now, as the Jenga tower topples, he says he enjoys the challenge of solving training tasks for robots —”It’s something new every day!”
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
AI算出
技術分析ainew評価標準
記事は Physical AI のデータ不足という課題に対し、脳波計測という新しいモダリティを導入する具体的な技術的解決策とその実証実験の詳細を報じており、単なるニュース報告を超えた技術的分析の価値がある。また、比較対象となる直前の記事が存在しないため、この独自のアプローチと実装知見は新規性が高いと判断される。
6つの評価軸を見る
- AI関連度
- 75
- 情報源の信頼性
- 75
- 新規性
- 75
- 調べる価値
- 50
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み