Grace Li 氏率いる新会社、AI ゲームの面白さを評価するツール「Design Arena」を発表
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
2025 年卒業直前に設立された Grace Li 氏の会社が、AI で生成したゲームの面白さを人間が判断する必要性から開発したスケール可能なフィードバック収集ツール「Design Arena」を公開した。
AI深層分析を開く2026年8月4日 23:14
AI深層分析
キーポイント
シードラウンドの成功と資金調達
Intelligence は Index Ventures の主導により、Conviction や A* などからの参加を得て 790 万ドルのシードラウンドを完了した。
人間による評価データの価値
自動ベンチマークがゲーム可能である一方、人間の判断は「面白さ」や「デザイン性」といった定性的な指標において不可欠であり、同社はこれをスケーラブルに提供している。
グローバルな嗜好データの蓄積
ユーザーがログインして出力を取得する仕組みにより、地域ごとのデザイン嗜好(例:アジアのウェブダッシュボードにおける極大主義的スタイル)や時間経過による変化を追跡可能である。
高収益と市場での地位確立
同社は現在 530 万人の利用者を抱え、年間収益(ARR)6000 万ドルを達成しており、AI 業界における人間主導の評価データの主要な供給源として位置づけられている。
ログインによるユーザー嗜好の追跡と自動化ベンチマークの限界
ユーザーがログインすることで、地域や時間経過に伴う嗜好の変化を追跡可能となる。これはハッキング事例で示されたように、大規模に実行できる自動化ベンチマークがゲーム化されるリスクを補完する重要な手段である。
重要な引用
"It was the missing bottleneck for a lot of these models to make improvements in the design space."
The site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry.
Li notes that web dashboards in Asia tend to have a more maximalist design style.
These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated
編集コメントを表示
編集コメント
「人間の味覚」を AI モデル開発のボトルネック解消に結びつけた事例は、生成 AI の成熟度が上がる中で不可欠な要素として認識されつつあることを示唆している。特に地域ごとの嗜好差をデータ化できる点は、グローバル展開する企業にとって戦略的に価値が高い。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
共同創業者のグレース・リー氏によれば、同社は 2025 年の卒業を数週間前に控え、数人の友人たちと協力して AI ゲームエンジンの開発に取り組むことから始まりました。生成されたモデルは機能するゲームを作れるものの、どれも面白くありませんでした。そこで「どうすればゲームが面白いと言えるのか?」という興味深い問いが生まれました。
人間による判断に代わるものはないと考えたチームは、大規模なスケールで誠実な人間のフィードバックを得る方法を模索し始めました。その結果生まれたのが Design Arena です。現在、この AI ツールは世界中で 530 万人に利用されています。実は多くの AI企業がスケーラブルなユーザーフィードバックを求めており、それに対して対価を支払う用意があることも判明しました。
「これは、これらのモデルがデザイン領域で改善を行うために欠けていたボトルネックでした」とリー氏は語ります。「約一週間後には、最先端の研究機関との最初の主要契約を締結し、その後は歴史が動いていきました。」
月曜日、Design Arena を開発する企業 Intelligence は、Index Ventures が主導し、Sarah Guo 氏と Mike Vernal 氏が率いる Conviction、A*、Valkyrie などが出資した 790 万ドルのシードラウンドを発表しました。
エンタープライズユーザーでない場合、Design Arena の利用は高度なモデルルーターを使うようなものです。プロンプトを入力するための ChatGPT 風のウィンドウがあり、ウェブサイトや画像など、12 種類以上の視覚フォーマットに対応する別々のドロップダウンメニューが用意されています。
リクエスト、フォーマット、スタイルを指定すると、「A vs. B」の選択形式で提示される出力を、最良のものから最悪のものまで順位付けするまで繰り返し行われます。
これは便利なサービスですが、プラットフォームの真価はエンタープライズ側にあります。参加するモデルは、メディア生成モデルに対する無限の即時フィードバック源としてこれを利用できます。ユーザーは自分が評価しているのがどのモデルかには無関心で、李氏(Li)が言うように「とにかく最良の結果を得たい」だけなので、その順位付けデータは、ユーザーが本当に求めているものを理解する上で決定的な入力となります。
最先端の研究ラボにとって、これは支払う価値のあるサービスだと李氏は述べています。同サイトは現在、年間約 6000 万ドルの ARR(年間経常収益)を発生させており、AI 業界における人間主導の評価データの主要ソースとしての地位を確固たるものにしています。
ただし、ユーザーは出力を受け取るためにログインする必要があり、これにより Intelligence は、異なる大陸や時間経過に伴う嗜好の変化を追跡することも可能になります(李氏によれば、アジアのウェブダッシュボードはよりマキシマリストなデザインスタイルを好む傾向があるそうです)。これらの対策は、自動化されたベンチマークの重要な補完手段です。自動化されたベンチマークは大規模に運用できますが、先週起きた Hugging Face の侵害事件 で劇的に示されたように、ゲーム化されたり操作されたりするリスクが常につきまといます。
ただし、クラウドソーシングによる人間のフィードバックが自動的に勝者市場になるとは限りません。Yupp は創業から 1 年未満で 今年初めに事業を停止しました。同社は a16z crypto のクリス・ディクソンから 3,300 万ドルを調達し、最先端モデルの一部を顧客として獲得し、ユーザー数も 130 万人を超えていたと主張していましたが、持続可能な長期的なビジネスを構築することはできませんでした。
それでも、人間の評価に基づく他のスタートアップは好調です。テキストベースの回答に類似のアプローチを採用する LM Arena は、有料製品の正式ローンからわずか 4 ヶ月後の今年 1 月に シリーズ A で 1 億 5,000 万ドルを調達しました。
当記事内のリンクを通じて購入が行われた場合、当社は小規模な手数料を受け取る可能性があります。これは当社の編集の独立性には影響しません。
ラッセル・ブランドは 2012 年以来テック業界を取材し続けており、プラットフォーム政策や新興技術に焦点を当てています。以前は The Verge や Rest of World で勤務しており、Wired、The Awl、MIT Technology Review などにも寄稿しています。問い合わせ先は russell.brandom@techcrunch.com または Signal(412-401-5489)です。
原文を表示
As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun?
There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale. The result became Design Arena, an AI tool now used by 5.3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
On Monday, the company behind Design Arena — dubbed Intelligence — announced a $7.9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others.
For non-enterprise users, using Design Arena is a lot like using a sophisticated model router. There’s a ChatGPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format, and style, you’ll be presented with a series of “A vs. B” choices until you’ve ranked the handful of outputs from best to worst.
It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want.
For frontier labs, that’s a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry.
Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time. (Li notes that web dashboards in Asia tend to have a more maximalist design style.) These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated, as the Hugging Face breach demonstrated in dramatic fashion last week.
That’s not to say that crowdsourced human feedback will be an automatic winning market. Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33 million from a16z crypto’s Chris Dixon. It too nabbed some frontier models as customers and had, it said, over 1.3 million users, but still couldn’t build a sustainable long-term business.
Even so, other startups based on human evaluation seem to be thriving. LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after formally launching its paid product.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.
He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み