視覚と言語モデルを用いた項目難易度予測
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
研究チームは、GPT-4.1-nanoを用いて、データ可視化リテラシーテスト項目の難易度を予測する手法を調査した。項目テキストと可視化画像の特徴を組み合わせ、米国成人の正答率を予測する能力を評価した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
arXiv:2603.04670v1 Announce Type: new
要約: 本研究は、データ可視化リテラシーテスト項目の難易度を判定する大規模言語モデル(LLMs)の能力を調査する。項目テキスト(質問と回答選択肢)、可視化画像、あるいはその両方から抽出した特徴量が、米国成人における項目難易度(正答率)を予測できるかどうかを検討する。GPT-4.1-nanoを用いて項目を分析し、これらの異なる特徴量セットに基づいて予測値を生成した。視覚特徴量とテキスト特徴量の両方を利用するマルチモーダルアプローチは、平均絶対誤差(MAE)が最も低く(0.224)、ユニモーダルな視覚のみのアプローチ(0.282)およびテキストのみのアプローチ(0.338)を上回った。性能が最も高かったマルチモーダルモデルを、外部評価用に確保していたテストセットに適用した結果、平均二乗誤差は0.10805となり、LLMsの心理測定分析および自動項目開発への応用可能性が示された。
原文を表示
arXiv:2603.04670v1 Announce Type: new
Abstract: This project investigates the capabilities of large language models (LLMs) to determine the difficulty of data visualization literacy test items. We explore whether features derived from item text (question and answer options), the visualization image, or a combination of both can predict item difficulty (proportion of correct responses) for U.S. adults. We use GPT-4.1-nano to analyze items and generate predictions based on these distinct feature sets. The multimodal approach, using both visual and text features, yields the lowest mean absolute error (MAE) (0.224), outperforming the unimodal vision-only (0.282) and text-only (0.338) approaches. The best-performing multimodal model was applied to a held-out test set for external evaluation and achieved a mean squared error of 0.10805, demonstrating the potential of LLMs for psychometric analysis and automated item development.
同じ出来事を2媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み