AI 利用実態、企業発表のみで独立検証源なしと研究者が指摘
本文の状態
日本語全文を表示中
詳細モードで約8分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MIT Technology Review AI
スタンフォード大学の研究者らが主導する「AI Observatory」が、大手企業による自己申告データにはない生きた利用実態を分析し、仕事以外の用途やリスク事象の割合が大幅に異なることを明らかにした。
AI深層分析を開く2026年8月18日 20:21
AI深層分析
キーポイント
独立検証機関の設立と目的
スタンフォード大学の STAIR ラボ所属アナ・ルエル氏らが主導する「AI Observatory」は、ユーザー同意のもと収集されたデータを用い、大手企業の自己申告に依存しない独立した AI 利用実態分析プラットフォームを構築した。
企業報告とのデータ乖離
Anthropic の経済インデックスなどの既存データを再分析した結果、同社が「仕事関連」としてフィルタリングしていた会話の約半数(48%)は、実際には健康、人間関係、成人向けコンテンツ、ハラスメントなど多岐にわたる非業務利用を含んでいた。
利用傾向とプラットフォーム応答の変化
2023 年から 2025 年にかけてのデータ分析により、プロンプトや会話ターン数が増加し、小話(small talk)の頻度も高まるなど、AI の使い方が時間とともに変化していることが示された。
モデルごとの利用傾向の差異
Grok はニュースや政治情報の検索に頻繁に使われるが誤情報も集中し、Anthropic はコーディング、Gemini はロールプレイ、ChatGPT は宿題支援にそれぞれ特化している。
同一モデル内でのバージョンによる違い
ユーザーは GPT-3.5 よりも GPT-4o との会話時間が長く、感情的な依存を生みやすい特徴があることが示された。
重要な引用
There is no independent source to corroborate it.
Stakeholders are currently making highly consequential decisions about AI's benefits and risks based on very limited data.
Nearly half of the conversations—or 48%—would have been filtered out.
"No single company report tells the whole story."
編集コメントを表示
編集コメント
企業発表のデータが「見せたい部分」だけを選別しているという指摘は、AI リスク管理において極めて重要な示唆を与える。業界全体が利用実態を正しく理解し、適切なガバナンスを構築するためには、こうした独立した第三者による検証データの重要性が今後さらに高まるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Anthropic や OpenAI といった AI 企業は、Claude や ChatGPT などの製品の利用状況について定期的にレポートを発表していますが、彼らが公開するのは自分たちにとって都合の良いデータだけだと、AI の研究者たちは指摘しています。
「それを裏付ける独立した情報源はどこにもありません」と、スタンフォード大学の信頼できる AI 研究(STAIR)ラボで博士課程を修了中のアンカ・ルエル氏は語ります。
ルエル氏はその新プロジェクト「AI オブザーバトリー」の共同リーダーを務めています。これは、Claude や Gemini といった人気モデルとの実際の会話データを、ユーザーの同意を得て収集された既存の 7 つのデータセットから集約・分析し、そのギャップを埋めることを目的とした公共プラットフォームです。このオブザーバトリーの狙いは、生成 AI の利用実態を評価する研究者や政策決定者に対して、独立した情報源を提供することにあります。ルエル氏によると、現在、ステークホルダーたちは極めて限られたデータに基づいて、AI の恩恵とリスクに関する重大な意思決定を下しています。
AI オブザーバトリーによる調査では、モデルごとの利用状況には大きな違いがあり、時間とともに変化していることが明らかになりました。また、主要 AI 企業のレポートで捉えられていない、より多くの機微な行動も検出されました。これらの企業は、個人の利用よりも業務利用に焦点を当てていると指摘されています。
「Anthropic Economic Index」は最も知名度が高く、広く引用される AI 利用データの一つですが、盲点もあります。その名が示す通り、この指標は Claude AI の業務や生産性に関連する利用に焦点を絞り、それ以外の会話は除外しています。
AI オブザーバトリーチームが Anthropic の手法を自社のデータセットに適用したところ、会話のほぼ半分(48%)がフィルタリング対象になると判明しました。フィルタリングされたのは仕事以外の会話で、そこには健康や人間関係に関する話題(Anthropic の分析では 31.2% に対し 44.2%)、成人向けまたは違法なトピック(7.9% vs 2.1%)、ハラスメントやヘイトスピーチ(27.5% vs 5.66%)、性的コンテンツ(16.7% vs 2.4%)が含まれる傾向が強く見られました。同様に、OpenAI の 2025 年 ChatGPT 利用に関する報告書でも、消費者の利用のわずか 30% が仕事関連であると指摘されています。
Anthropic は、Claude をサポートや伴侶として、あるいは児童性的虐待材料(CSAM)の生成に利用する事例について個別のブログ記事を公開しています。しかし、AI オブザーバトリーの「鳥瞰図」的な分析は、別々の報告書に分割されるよりも、研究者が AI の多様な利用方法をより一貫して理解する助けになると、UT-Austin 情報学部で AI システムとの人間関係を探求しているデイビッド・ウィダー准教授(AI オブザーバトリーには関与していない)は述べています。
AI オブザーバトリーが調査したデータセットには 2023 年から 2025 年にかけての会話が含まれており、人々が AI をどのように利用しているか、そして各 AI プラットフォームがどう対応しているかに違いがあることが明らかになりました。
AI オブザーバトリーの研究に含まれる最も大規模で詳細なデータセットの一つである WildChat では、時間の経過とともにプロンプトトークン数、レスポンストークン数、会話のターン数が増加し、会話がより長くなり、内容も洗練されていく傾向が示されました。
また、時間の経過とともに雑談の頻度も顕著に増加しました。これは AI との交流が深まっていることを示唆しており、同時に AI アシスタント側から「自分はチャットボットです」といった自己開示が減っていることもわかります。
さらに、研究者らが「機密性の高い利用」と分類したやり取り、つまり性的ハラスメントやヘイトスピーチなど有害または制限される可能性のあるコンテンツを含むものは減少しました。これはプラットフォーム側がより効果的な安全対策を講じている可能性を示しています。
AI オブザーバトリーによる調査では、モデルによって AI の利用様式に明確な違いがあることも判明しました。ツールごとに、ユーザーが扱う話題や対話スタイル、会話の構造、そして機密性の高い利用事例が発生する確率とその種類まで多岐にわたります。
例えば、研究者らは Grok と Gemini が情報検索のために頻繁に使われていることを発見しました。特に Grok はニュースや政治に関する情報の取得で人気がありましたが、一方で誤情報が集中しやすい場所でもありました(これは、Grok 上で誤情報が容易に拡散する様子を示す他の研究とも一致しています。xAI はコメント依頼に対して回答しませんでした)。
一方、コーディングには Anthropic を、社会的な役割やロールプレイには Gemini を、宿題のサポートには ChatGPT を利用する傾向が強いことがわかりました。
同じモデルのバージョン間でも、利用状況には違いが見られました。研究者らは、ChatGPT が GPT-3.5 で動作しているときは会話が短く、GPT-4o の場合はより長く、反復的なものになることを発見しました。これは、後者のバージョンが「感情的な依存」を生みやすいことで知られているため、納得できる結果です。
しかし、企業の報告書では、自社モデル間や異なるモデル間のこうした微妙な違いを捉えきれていないのが実情です。「どの企業の報告書も、全体像を単独で語ることはできません」と語るのは、MIT メディア・ラボから最近博士号を取得した Shayne Longpre 氏です。同氏は Reuel と共同でこの研究を主導しました。
AI オブザーバトリーを構築するため、Reuel 氏と MIT、スタンフォード大学、データ・プロベナンス・イニシアチブなど複数の機関の研究者たちは、過去の研究によって収集された 7 つの実世界データセットから、24,521 の会話(計 85,633 の対話ターン。つまり、ユーザーからのプロンプトとそれに対する AI の応答のペア)を統合しました。これら 2023 年から 2025 年にかけての会話は、ChatGPT、Gemini、Claude、Grok など 52 の異なるモデルを利用した 5,000 人のユーザーによるものです。
ただし、これらの会話データは、大手ラボが実際に保有している膨大なデータに比べれば、ほんのわずかな一部に過ぎません。例えば、最新の Anthropic Economic AI Index は 100 万件の Claude の会話分析に基づいていますし、OpenAI が ChatGPT の利用実態について発表した報告書も、150 万件の会話を分析したものです。
Anthropic の代表者は、同社が発表した研究は、研究チームが抱く特定の疑問や関心を反映したものであり、外部の独立系研究を支援することの重要性を示すものだと説明しました。OpenAI はコメント依頼に対して回答していません。
さらに、AI オブザーバトリーのデータセットが自発的に提供された情報に基づいているという事実は、機密性の高い利用事例が過小評価されている可能性を示唆しています。こうした利用は、ユーザーが共有しようとしにくい傾向があるためです。そのため、研究者たちは今回の調査結果が AI 利用のすべてを代表するものではないと注意を促しています。
しかし、オブザーバトリーの取り組みは研究コミュニティへのアクセスを広げる点で意義があります。AI 企業は通常、分析のためにチャットデータを共有しません。その結果、企業の報告書は自社を好意的に映す発見に焦点が当たりがちです。Reuel 氏や Widder 氏のような独立系の研究者たちはこの点を指摘しています。
「例えば、『Anthropic の汎用 AI システムは主に善のために使われているのか、それとも悪のために使われているのか』といった問いに対して答えられる方法がないのが現状です」と UT オースティン校情報学部准教授の Widder 氏は説明します。「その情報は企業秘密だからです」
AI Observatory のデータは研究者が分析に利用できるようになり、チームは今後データの拡充も目指しています。理想を言えば、Reuel 氏は AI 企業がユーザーのプライバシーを守りつつ、独立した研究者とデータを共有することを望んでいます。しかし現状では、AI の利用データに基づいて意思決定を行うことは、「企業の語る物語の背後にある実態が見えないまま、荒れた野外で完全に活動し、重大な決断を下すリスクを冒すこと」と同義だと Reuel 氏は指摘しています。
原文を表示
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say.
“There is no independent source to corroborate it,” says Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab.
Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill in the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The Observatory’s intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI. Stakeholders are currently making highly consequential decisions about AI’s benefits and risks based on very limited data, says Reuel.
The AI Observatory found that AI use differs significantly across models, and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use.
Anthropic Economic Index is one of the best known and most widely cited sources of AI usage data but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses.
When the AI Observatory team applied Anthropic’s methods to their dataset, they found that nearly half of the conversations—or 48%— would have been filtered out. Those non-work-related conversations that were filtered out were more likely to include health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT use, similarly, found that only 30% of consumer use was related to work.)
Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, “but having [the AI Observatory’s] bird eye view analysis rather than sectioned off into a separate report helps” researchers understand the different uses more consistently, says David Widder, an assistant professor at UT-Austin’s School of Information, who researches how people interact with AI systems and is not involved with the AI Observatory.
The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded.
Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, indicated by increases in prompt tokens, response tokens, and conversation turns.
There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile the AI assistants’ self-disclosure (i.e. that it was a chatbot) decreased.
Additionally, exchanges that the researchers labeled as sensitive use—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—dropped. That might suggest that platforms were generally deploying more effective safeguards.
The AI Observatory also found that AI use looked significantly different depending on the model. Depending on the tool, users ranged in topics, interaction styles, conversation structures, as well as both the likelihood and type of sensitive use cases.
For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has also shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.)
Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.
There were even differences among different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction.
Company reports, however, didn’t tend to capture these nuances across or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel.
To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions, aggregated 24,521 conservations across 85,633 conversational turns (that is, the user prompt and corresponding AI response) from seven real-world datasets collected by previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.
But these conversations are a drop in the proverbial bucket compared to the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.
An Anthropic representative said that their published research reflects their research teams’ specific questions and interests, and the importance of supporting external independent research. OpenAI did not respond to requests for comment.
Additionally, the fact that the AI Observatory’s dataset draws from voluntarily-provided sources means that it’s likely underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use.
The Observatory’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint their companies in the best light, independent researchers, like Reuel and Widder, say.
“When we want to ask, for example: is Anthropic’s general-purpose AI system…used mostly for good, or mostly for bad…we don’t have a way of answering that question because that information is proprietary,” explains Widder, the assistant professor at UT-Austin’s School of Information.
The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives,” says Reuel.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み