DesignArena のクリエイターたち、AI モデルの審美性を向上させるため 790 万ドルを調達
本文の状態
日本語全文を表示中
詳細モードで約4分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
デザインプラットフォーム「DesignArena」のクリエイターらが、AI モデルの品質や審美性を高めるための資金として 790 万ドル(約 12 億円)を調達したと報告された。
AI深層分析を開く2026年8月4日 05:12
AI深層分析
キーポイント
人間による評価の重要性と市場価値
AI モデルが機能的であっても「面白さ」を判断するには人間の審美眼が必要であり、同社はこれをスケーラブルに提供することで業界のボトルネックを解消した。
強力な財務実績と資金調達
同社は現在年間6000万ドル(ARR)の収益を上げており、この実績に基づき Index Ventures 主導で790万ドルのシードラウンドを完了した。
自動化ベンチマークの限界と代替
Hugging Face の侵害事例が示すように自動評価は改ざんのリスクがあるため、ログインユーザーの行動データに基づく人間主導の評価が不可欠な補完手段となっている。
グローバルな嗜好データの蓄積
ユーザーのログインデータを分析することで、地域や時間経過に伴うデザイン嗜好の変化(例:アジアにおけるマキシマリスト傾向)を把握できる独自の価値を提供している。
Crowdsourced human feedback market risks
Yupp は1年以内に閉鎖され、持続可能なビジネスを構築できなかった事例がある。
重要な引用
"It was the missing bottleneck for a lot of these models to make improvements in the design space."
The site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry.
These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated.
Yupp shuttered its doors earlier this year after raising $33M from a16z crypto's Chris Dixon.
編集コメントを表示
編集コメント
AI モデルの性能評価において、数値的なベンチマークだけでなく人間の主観的・文化的な「面白さ」やデザイン嗜好を定量化する市場が急成長している。このニュースは、開発者がモデルの完成度を高めるために、いかにして信頼性の高い人間フィードバックをスケーラブルに獲得するかという課題への具体的な解決策を示すものである。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
共同創業者のグレイス・リーによれば、同社は2025年の卒業の数週間前に設立されました。当時、数人の友人が協力してAIゲームエンジンの動作を試みていたのです。
モデルは機能的なゲームを生成できましたが、どれも面白くありませんでした。そこで「どうすればゲームの面白さを判断できるのか?」という興味深い問いが生まれました。
彼らは人間の判断に代わるものはないと結論づけ、すぐに大規模な人間からの正直なフィードバックを得る方法を模索し始めました。
その結果生まれたのが DesignArena です。現在、このAIツールは世界中で530万人に利用されています。実は多くのAI企業がスケーラブルなユーザーフィードバックを求めており、それに対して対価を支払う用意があることも判明しました。
「これは、これらのモデルがデザイン分野で改善を行うために欠けていたボトルネックでした」とリー氏は語ります。「約1週間後には最先端ラボとの最初の主要契約を結び、その後は歴史の通りです」
月曜日、DesignArenaの開発元である Intelligence は、Index Venturesが主導し、Conviction(Sarah Guo氏とMike Vernal氏)、A*、Valkyrieなどが出資する790万ドルのシードラウンドを完了したと発表しました。
エンタープライズユーザー以外にとって、DesignArena の利用は高度なモデルルーターを使うようなものです。プロンプト用の Chat-GPT 風のウィンドウがあり、ウェブサイトや画像、その他 dozen 種類のビジュアルフォーマットを切り替えるためのドロップダウンメニューが用意されています。リクエスト、フォーマット、スタイルを入力すると、「A vs. B」の選択肢が次々と提示され、最終的に数ある出力結果を最良のものから最悪のものへと順位付けする仕組みになっています。
これは有用なサービスですが、プラットフォームの真価はエンタープライズ側にあります。参加するモデルにとっては、メディア生成モデルに対する絶え間ない即座のフィードバック源として機能するからです。ユーザーは自分が評価しているのがどのモデルかには無関心で、李氏(Li)が言うように「とにかく最良の結果を得たい」というだけなので、その順位付けデータこそが、ユーザーが本当に求めているものを理解するための決定的な入力となります。
フロンティアラボにとってこれは有料に値するサービスだと李氏は語ります。同サイトは現在、年間収益(ARR)で 6,000 万ドルを達成しており、AI 業界における人間主導の評価データの主要ソースとしての地位を確固たるものにしています。
重要なのは、ユーザーが出力を受け取るためにログインする必要がある点です。これにより Intelligence は、異なる大陸や時間経過とともに「好み」がどのように変化するかを追跡できます。李氏によると、アジアのウェブダッシュボードはよりマキシマリストなデザインスタイルを持つ傾向があるそうです。
これらの対策は、自動化されたベンチマークを補完する重要な要素です。自動化されたベンチマークは大規模に運用可能ですが、先週起きた Hugging Face の侵害事件 が示したように、ゲーム化されたり操作されたりするリスクが常にあります。
ただし、クラウドソーシングによる人間のフィードバックが自動的に市場で勝つことを意味するわけではありません。創業から 1 年未満の Yupp は、今年初めに閉鎖されました。同社は a16z crypto のクリス・ディクソン氏から 3,300 万ドルを調達し、最先端モデルを顧客として獲得し、ユーザー数も 130 万人を超えていましたが、持続可能な長期的なビジネスを構築することはできませんでした。
それでも、人間の評価に基づく他のスタートアップは thriving(繁栄)しています。テキストベースの回答に類似のアプローチを採用する LM Arena は、有料製品の正式ローンからわずか 4 ヶ月後の今年 1 月に シリーズ A で 1.5 億ドルを調達 しました。
当記事内のリンクを通じて購入が行われた場合、当社は少額のコミッションを獲得する可能性があります。これは当社の編集の独立性には影響しません。
ラッセル・ブランドは 2012 年以来テック業界を取材しており、プラットフォームポリシーや新興技術に焦点を当てています。以前は The Verge や Rest of World で勤務し、Wired、The Awl、MIT Technology Review にも寄稿しています。
原文を表示
As co-founder Grace Li tells it, her company started a few weeks before graduation in 2025, with a handful of college friends trying to make their AI game engine work. The models could make functional games, but none of the games were fun — which raised the interesting question, how can you tell if a game will be fun?
There was no substitute for human judgment, they decided, and soon they were brainstorming ways to get honest human feedback at scale.
The result became DesignArena, an AI tool now used by 5.3 million people around the world. As it turned out, there were lots of AI companies looking for scalable user feedback — and many of them were willing to pay for it.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
On Monday, the company behind DesignArena — dubbed Intelligence — announced a $7.9 million seed round led by Index Ventures with participation from Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others.
For non-enterprise users, using DesignArena is a lot like using a sophisticated model router. There’s a Chat-GPT-style window for prompts, with separate dropdowns for websites, images, and a dozen other visual formats. Once you put in the request, format and style, you’ll be presented with a series of “A vs. B” choices until you’ve ranked the handful of outputs from best to worst.
It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want.
For frontier labs, that’s a service worth paying for, Li says, adding the site is currently generating $60 million in ARR, solidifying its position as a key source of human-led evaluation data for the AI industry.
Crucially, users have to log in to get their output, so Intelligence can also track how those tastes change across different continents and over time. (Li notes that web dashboards in Asia tend to have a more maximalist design style.) These measures are an important complement to automated benchmarks, which can operate at a greater scale but are often subject to being gamed or otherwise manipulated, as the Hugging Face breach demonstrated in dramatic fashion last week.
That’s not to say that crowdsourced human feedback will be an automatic winning market. Less than a year after launching, Yupp shuttered its doors earlier this year after raising $33M from a16z crypto’s Chris Dixon. It too nabbed some frontier models as customers and had, it said, over 1.3 million users, but still couldn’t build a sustainable long-term business.
Even so, other startups based on human evaluation seem to be thriving. LM Arena, which takes a similar approach to text-based responses, raised $150 million in a Series A in January, just four months after formally launching its paid product.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT’s Technology Review.
He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489.
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み