GPT-6 Astra
GPT-6 Astra
コメントの核心
ベンチマークと AGI 判定
意見が分かれるARC-AGI-3 のスコア解釈を巡り、レスポンス API ハーネスの影響や「Fable」との比較から、実質的な AGI 到達の可能性が議論されています。一方で、ベンチマークの過熱した期待に対する懐疑論も存在します。
代表コメント(原文)
“The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would…”
リリース形態と業界の方向性
意見が分かれる動画や記者会見のない地味な発表形式に対し、契約上の圧力によるものではないかという推測が示されました。他方、プロジェクトを整理した結果として OpenAI が本質に回帰し、競合他社との差を縮めたとする評価も確認できます。
代表コメント(原文)
“I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The…”
全論点と英語の原コメントを見る
議論の全体像
GPT-6 Astra の発表を巡り、ARC-AGI-3 ベンチマークのスコア解釈や「AGI」到達の定義を巡る議論が活発です。一方で、大規模な発表会がない静かなリリース形態への違和感や、モデル進化のスピードに疲弊する声も上がっています。
ベンチマークと AGI 判定
ARC-AGI-3 のスコア解釈を巡り、レスポンス API ハーネスの影響や「Fable」との比較から、実質的な AGI 到達の可能性が議論されています。一方で、ベンチマークの過熱した期待に対する懐疑論も存在します。
“The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the…”
“I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547 Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more…”