AI ラボはペリカンテスト対策か?
TLDR AI は、AI ラボが「Pelicanmaxxing」と呼ばれるベンチマーク操作を行っている可能性を検証するため、7 つの最先端モデルを対象に独自の SVG 生成実験を行い、その結果と分析手法を公開した。
AIニュース価値スコアβ
技術分析AI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 50
- 新規性
- 75
- 検索具体性
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
AI モデルのベンチマーク対策に関する疑念に対し、著者が独自に大規模な生成実験を行い、具体的なスコアと結果を提示した点で高い技術的価値を持つ。また、GPT-5.6 や Claude Sonnet 5 など最新モデル名が明記されており検索機会も高い。
キーポイント
Pelicanmaxxing の検証実験
著者は Simon Willison 氏の有名な「ペリカンが自転車に乗る」プロンプトを拡張し、8 種類の動物と 6 種類の車両の組み合わせ(計 48 プロンプト)を用いて 1,008 個の SVG を生成・評価した。
LLM ジャッジによる定量化
生成された画像を人間ではなく LLM ジャッジ(Claude Fable 5)に評価させることで、各モデルの視覚的生成能力を客観的にスコアリングし、ベンチマーク操作の有無を検証した。
業界全体への示唆
数十億ドル規模の投資とユーザー獲得競争の中で、特定のベンチマークで結果を最適化する「Pelicanmaxxing」のインセンティブが実際に働いているかという懸念に対し、データに基づく分析を提供した。
実験の規模と手法
著者は7つの主要モデルを対象に、8種類の動物×6種類の車両の組み合わせ(計48プロンプト)で1,008枚のSVG画像を生成し、LLMによる多段階評価を行いました。
ベンチマーク最適化の仮説検証
AI ラボが「ペリカン×自転車」のベンチマークに過剰適合(pelicanmaxxing)している場合、この特定の組み合わせや関連する要素で他モデルより著しく高いスコアを示すはずだと仮定しました。
視覚的評価の結果
スコアリング前の画像を直接確認した結果、どのモデルも「ペリカンが自転車に乗る」描写において他モデルよりも特段に優れているようには見えず、最適化の明確な証拠は見つかりませんでした。
主観的評価の限界
著者は画像を直接確認したが、ペリカンと自転車の組み合わせが他のモデルよりも際立って優れているとは感じられず、特定の動物-車両組み合わせの良し悪しがモデル全体の能力と相関している可能性を示唆した。
重要な引用
"Generate an SVG of a pelican riding a bicycle"
"When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn't it be tempting to pelicanmax your model just a bit?"
"I built a grid of 8 animals × 6 vehicles = 48 prompts"
"The benchmark is now famous enough that there's plenty of discussion about its usefulness and about whether AI labs might be benchmaxxing on it."
"My hypothesis is that if a lab trained on the benchmark, it should show up in some combination of the pelican row scoring above what the animal deserves, the bicycle column scoring above what the vehicle deserves, or the specific pelican-bicycle cell beating both."
Otherwise they look like the rest of what each model draws, and the labs that draw good pelicans on bicycles also do a good job drawing other animal-vehicle combinations.
影響分析・編集コメントを表示
影響分析
この記事は、AI モデルの性能評価において「ベンチマーク・ゲーミング」が実際に発生している可能性を浮き彫りにし、業界関係者や投資家に対して、公式ベンチマーク以外の非公式指標(特に視覚的生成能力)の重要性と限界を再考させるきっかけとなる。また、LLM を用いた自動評価手法の実証を通じて、今後のモデル比較における新しいアプローチの可能性を示唆している。
編集コメント
「Pelicanmaxxing」という造語が示す通り、AI モデルの評価基準は数値だけでなく、コミュニティが共有する文化的・視覚的指標にも左右される複雑な構造を持っています。この分析は、ベンチマークの信頼性を再検証する上で重要な視点を提供しています。
ここ数年、Simon Willison は主要な大規模言語モデル(LLM)のリリースごとに、同じプロンプトを使ってテストを行ってきました。そのプロンプトとは「自転車に乗るペリカンの SVG を生成せよ」というものです。
当初は皮肉めいたベンチマークとして始まったこの試みは、AI 業界における最も有名なインフォーマルな評価基準の一つとなりました。Simon が作成した「自転車に乗るペリカン」の画像は、AI ラボが新モデルを発表する際の Hacker News のスレッドにおいて、最も高評価を集めるコメントの常連となっています。
このベンチマークはすでに有名になりすぎており、その有用性や、AI ラボがこのテストでベンチマックス(ベンチマークを最適化すること)を行っているのではないかという議論が多数 あります。さらに、数十億ドル、あるいは兆単位の資金が絡む状況で、強力な結果がユーザーの信頼を得る助けになるなら、モデルを少しだけ「ペリカンマックス(ベンチマーク向けに最適化すること)」したくなる誘惑は確かにあるでしょう。
私はその真相を知りたくて、小さな実験を行いました。最先端の 7 つのモデルに対して合計 1,008 枚の SVG を生成し、LLM による判定で採点しました。分析には Claude Fable 5 を使用しています。
この記事ではその結果を報告します。すべてのコードは Github で公開されています。
テスト方法
私は、8 種類の動物と 6 種類の乗り物を組み合わせたグリッドを作成しました。これは合計 48 のプロンプトとなり、その中に「自転車に乗るペリカン」という有名なプロンプトも含まれています。
対象とした動物はペリカン、フラミンゴ、アオサギ、カワウソ、アライグマ、アンテロープ、クジラ、そして猫です。
乗り物としては、自転車、一輪車、スケートボード、スクーター、飛行機、ボートを設定しました。
すべてのプロンプトはサイモン氏のものをほぼそのまま流用し、動物と乗り物の名前だけを変更しています。動物と乗り物の選定には厳密な手法を用いたわけではありませんが、元のプロンプトとの類似度や難易度のバランスを考慮して多様な組み合わせを試みました。フラミンゴとアオサギはペリカンに似ており、猫、アライグマ、カワウソは比較的容易なケースです。一方、アンテロープは難しく、クジラは最も異なる存在と言えます。
OpenRouter を通じて 7 つのモデルをテストしました。対象は「GPT-5.6 Terra」、「Claude Sonnet 5」、「Gemini 3.5 Flash」、「Grok 4.5」、「Qwen3.7-Max」、「GLM-5.2」、そして「DeepSeek V4 Pro」です。各プロンプトに対して温度パラメータを 1.0 に設定し、すべてのモデルに同程度の推論負荷がかかるよう要求してサンプルを 3 件ずつ生成しました。その結果、合計 1,008 個の SVG ファイルが作成されました。
次に、生成された画像を 3 つの工程からなるパイプラインに通しました。
- レンダリング: 各 SVG を PNG 形式に変換します。モデルが SVG を返さない場合や、レンダリングに失敗した場合は有効なファイルが得られるまで再生成し、その試行回数を記録します。1,008 件の生成全体で再試行が必要だったのはわずか 11 回でした。
- 判定: GPT-5.6 Luna が各画像に対して、動物の描写、乗り物の描写、そして動作の一貫性の 3 つについてそれぞれ 1 から 5 の評価を行います。以下で動物や乗り物をランク付けする際は、それぞれの項目に付与された評価値をそのまま使用します。一方、画像ごとに単一の数値が必要となる場合は、これら 3 つの評価の平均値を採用し、「判定スコア」と呼ぶことにしました。
特徴抽出:より詳細な分析のため、レンダリングされた各画像を Gemini 3.1 Flash-Lite に通し、認識した動物や車両、被写体の向き、そしてシーン要素のオープンエンドなリストを記録させました。
私の仮説は、もしあるラボがベンチマーク用に訓練されていたなら、その結果として「ペリカン」行で動物本来の評価以上、「自転車」列で車両本来の評価以上、あるいは「ペリカン×自転車」という特定のセルで両者を上回るスコアを示すはずです。
根拠1:ペリカンの乗った自転車は、見た目も良くならない
スコアリングを行う前に、最も単純なテストとして画像自体を直接確認しましょう。あるラボが描いたすべての画像と、各画像下の判定スコア(クリックで拡大表示)を見てみます。
| Lab | Sample | pelican | bicycle | failed |
|---|---|---|---|---|
| A4 · V5 · C4 | unicycle | failed | A5 · V4 · C5 | |
| skateboard | failed | A5 · V4 · C5 | ||
| scooter | failed | A4 · V4 · C3 | ||
| plane | failed | A4 · V5 · C5 | ||
| boat | failed | A5 · V5 · C5 |
| Lab | Sample | flamingo | bicycle | failed |
|---|---|---|---|---|
| unicycle | failed | A4 · V4 · C4 | ||
| skateboard | failed | A4 · V5 · C3 | ||
| scooter | failed | A5 · V5 · C5 | ||
| plane | failed | A4 · V5 · C3 | ||
| boat | failed | A4 · V5 · C4 |
| Lab | Sample | heron | bicycle | failed |
|---|---|---|---|---|
| unicycle | failed | A4 · V4 · C5 | ||
| skateboard | failed | A4 · V5 · C5 | ||
| scooter | failed | A5 · V5 · C5 | ||
| plane | failed | A5 · V1 · C1 | ||
| boat | failed | A5 · V5 · C5 |
| Lab | Sample | otter | bicycle | failed |
|---|---|---|---|---|
| unicycle | failed | A5 · V3 · C3 | ||
| skateboard | failed | A5 · V5 · C5 | ||
| scooter | failed | A5 · V4 · C5 | ||
| plane | failed | A3 · V5 · C5 |
boat
failed
A4 · V5 · C4
raccoon
bicycle
failed
A5 · V3 · C4
unicycle
failed
A5 · V3 · C3
skateboard
failed
A5 · V5 · C5
scooter
failed
A5 · V5 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
antelope
bicycle
failed
A5 · V5 · C3
unicycle
failed
A4 · V4 · C4
skateboard
failed
A4 · V5 · C4
scooter
failed
A4 · V5 · C3
plane
failed
A5 · V5 · C4
boat
failed
A5 · V5 · C5
whale
bicycle
failed
A5 · V3 · C3
unicycle
failed
A5 · V3 · C3
skateboard
failed
A5 · V4 · C5
scooter
failed
A5 · V5 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
cat
bicycle
failed
A5 · V3 · C3
unicycle
failed
A5 · V4 · C3
skateboard
failed
A5 · V5 · C5
scooter
failed
A5 · V4 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
私は以下の分析を実行する前に、画像を自分で確認しました。特に目立つ点は見つかりませんでした。
ペリカンと自転車の組み合わせが、そのモデルのグリッド全体の中で際立って優れているケースは見当たりませんでした。GLM-5.2 の最初のサンプルでは、他のものよりわずかに良く見えた気もしますが、そのバッチにはスケートボードに乗るカッコいいアオサギも生成されており、断定はできません。
それ以外は、各モデルが描く他の画像と大差ありません。ペリカンの自転車を描くのが上手なラボは、動物と車両の組み合わせを他のものでも同様に上手に描いています。
しかし、このテストは再現が難しく、人によって評価も異なります。そのため、より定量的な指標が必要だと考え、上記の方法を採用しました。
証拠その2:ラボはペリカンの描画において他社よりも優れていない
全モデルを統合した動物ごとの平均評価点は以下の通りです。
ペリカンは8種中6位で、猫、クジラ、アライグマ、サギ、アンテロープに次いでいます。もし AI ラボがベンチマークデータを用いて学習を行っていたのであれば、ペリカンが上位に来るはずです。しかし実際には、下位半分にとどまっています。7 つのラボすべてにおいて、猫やクジラ、アライグマの描画の方がペリカンのそれよりも優れています。
もちろん、ペリカンは猫に比べて描くのが難しいだけかもしれません。ペリカンに特化して学習を行っても、簡単な動物には及ばないままになる可能性もあり、この順位付けだけではその可能性を否定できません。難易度による影響については、証拠その4 で補正を行います。
証拠その3:ラボは自転車の描画においても他社よりも優れていない
自転車はさらに評価が低く、最下位に近い飛行機とほぼ同点で2番目に低い順位です。
もし AI ラボがベンチマークデータを用いて学習を行っていたのであれば、自転車はこのランキングの上位に来るはずです。しかし実際にはそうではありません。ただし、前述と同様の注意が必要です。スケートボードに比べて自転車の描画は複雑です。両端の車軸をつなぐフレーム、一対の車輪、ハンドル、シート、ペダルなどが必要であり、これらが欠落していたり接続が不十分だったりすると、評価者は自転車画像の約 3 分の 2 で減点します。自転車画像で学習を行っても、より単純な乗り物と比較すれば完璧な描画ができない可能性は残ります。
飛行機に関する一つの補足ですが、"airplane"ではなく"plane"を選んだのは失敗でした。モデルはこれを幾何学的な形状として解釈する傾向があるためです。実際、描かれたのは飛行中の航空機ではなく、平らな地面に立つ動物でした。
"plane"の画像では、他の5種類の乗り物(ゼロ件)と比較して、特徴抽出器が車両を全く検出できないケースが25件(全168件中)ありました。また、車両としての評価が1または2点だった割合も20%に達し、これは自転車(5%)、ボートやスクーター、スケートボード(いずれもゼロ)と比べて顕著な差です。
証拠4:実験室は、自転車のペリカンの描画においても他より優れていない
これらを組み合わせると、「自転車のペリカン」のスコアは48項目中42位という最下位付近に位置します。
しかし、組み合わせによっては単に描画が難しいものもあるかもしれません。これを考慮するため、1,008枚の画像すべてに対して固定効果回帰分析を行いました。モデルは「スコア ~ 実験室 + 動物×乗り物」とし、ペリカン、自転車、およびその組み合わせ(ペリカン×自転車)に関する各実験室ごとの交互作用項を含めました。標準誤差には堅牢な推定量を使用しています。
ここでいう"animal × vehicle"の項は、48種類の組み合わせそれぞれが持つ固有の難易度を吸収します。一方、交互作用項は、各実験室が基準となる組み合わせに対して平均的な実験室よりもどれだけスコアを押し上げているか(または引き下げているか)を、信頼区間付きで測定するものです。
その結果は以下の通りです。
- 各実験室におけるペリカンの効果(6種類の乗り物すべてにわたるペリカンへのブースト)は、判定スコアで-0.11から+0.14の範囲内に収まり、いずれも統計的に有意な差ではありませんでした。最も小さいp値でも0.25です。
各ラボごとの自転車効果(8 種類の動物すべてにおける自転車の効果に対するラボの寄与)は、Grok 4.5 の -0.18 (p=0.11) から Gemini 3.5 Flash の +0.27 (p=0.022) まで様々です。統計的に有意水準 p < 0.05 をクリアしたのは Gemini だけですが、他の 7 つは正負両方向にばらついています。
「ペリカンと自転車」の組み合わせ特有の効果(各効果の単純な合計を超えた追加的なブースト)も、p < 0.05 をクリアしていません。最も大きな正の値を示したのは GLM-5.2 の +0.35 (p=0.12) ですが、これは先ほど言及したものです。今回の実験において信号に近いものとしてはこれが最良ですが、依然として偶然の範囲内です。
以下が各ラボごとの推計値一覧です。ペリカンマックス(結果操作)を行っているラボであれば、その行全体でゼロラインより右側にプロットされた点が並ぶはずです:
すべてのペリカンの区間とすべてのセルの区間にゼロが含まれています。例外はただ一つだけあります。それは自転車列における *Gemini 3.5 Flash* です。しかし、21 回の検定を行って p < 0.05 を基準とする場合、偶然だけで約 1 つの偽陽性(21 × 0.05 ≈ 1.05)が現れることが予測されます。実際に出現した数もまさに 1 つです。また、これは多重比較補正にも耐えられません。21 回の検定に対するボンフェローニ閾値は 0.05/21 ≒ 0.002 ですが、Gemini の p 値は 0.022 です。推計値と p 値の全表は リポジトリ にあります。
ただし、これらの区間は幅広で、平均して約±0.6 の判定者ポイントに相当します。これより小さなブーストはこのテストでは検出できません。
証拠 #5:ペリカンと自転車のシーンは記憶されていない
ペリカンが自転車に乗っている画像は、記憶された構成のように見えるという指摘があります。具体的には、ペリカンが常に右向きであることや、太陽やマフラーといった共通の要素が繰り返し登場するパターンなどが根拠として挙げられています。そこで、この説が事実かどうかを検証してみました。
方向性: 7 つの研究機関が生み出した全 21 枚の「ペリカン×自転車」画像は、すべて右向きです。他の動物と車両の組み合わせでこれほど一貫して同じ方向を向く例はありません。
ただし、右向きであること自体は珍しくありません。調査対象となった 1,008 枚の画像のうち 60% が右向きでした。その頻度は動物や車両の種類によって異なり、自転車はその傾向が最も強い車両の一つです。
| 車両 | 左向き | 右向き | 不明 |
|---|---|---|---|
| スクーター | 9% | 83% | 8% |
| 自転車 | 11% | 81% | 8% |
| スケートボード | 19% | 60% | 21% |
| プレイン(飛行機) | 21% | 58% | 21% |
| ユニサイクル | 22% | 45% | 33% |
| ボート | 33% | 35% | 32% |
ペリカンもまた、右向きになりやすい動物のリストに含まれています。
| 動物 | 左向き | 右向き | 不明 |
|---|---|---|---|
| アンテロープ | 21% | 78% | 1% |
| ペリカン | 22% | 78% | 0% |
| コウノトリ | 22% | 77% | 1% |
| クジラ | 34% | 65% | 1% |
| フラミンゴ | 36% | 64% | 0% |
| アット(アザラシ) | 8% | 45% | 47% |
| ネコ | 8% | 40% | 52% |
| アライグマ | 3% | 36% | 61% |
ペリカンや自転車を正面から描こうとすると難しく、モデルはほぼ例外なく横顔(左向きか右向き)で描きます。そのため、方向が不明確な画像が極めて少ないのです。他の組み合わせでも同様の傾向が見られます。例えば「アンテロープ×スクーター」では 21 枚中 20 枚、「コウノトリ×自転車」では 21 枚中 19 枚が右向きでした。つまり、21 枚すべてが右向きであることは、統計的に外れ値(アウトライヤー)とは言い難いのです。
シーン要素: 抽出ツールに画像内の要素を自由に名付けさせたところ、以下の出現頻度が確認されました:
記憶されたシーンとは、同じ要素の組み合わせが画像を繰り返すたびに現れることを指します。私はその痕跡を探し、いくつかの組み合わせでは毎回同じ要素が出現する傾向があることを見つけました。例えば、ボートに乗ったフラミンゴには必ず太陽が含まれています。飛行機に乗ったカワウソは 38% の確率でスカーフを身につけており、自転車に乗った猫も 38% の確率でバスケットが付いています。
一方、自転車のペリカンには特に目立った特徴は見当たりません。他の動物と車両の組み合わせと同様に、特定の要素が頻繁に現れるだけなのです。
限界点
- LLM 判定器を単一に使用している点: ここでのスコアリングはすべて、GPT-5.6 Luna という一つのモデルが画像を一枚ずつ評価した結果です。自己整合性の確認や調整にはあまり時間を割いておらず、再実行時に同じ判断を下す頻度も検証していません。もしこのモデルが絵の質を信頼できる形で判定できないのであれば、上記の数値に大きな意味はありません。また、この判定器は参加者の一つである GPT-5.6 Terra と同じファミリーに属しています。ただし、すべてのラボが 48 通りの組み合わせを描画するため、特定のラボのスタイルを好む判定器が現れたとしても、その影響は一度に全体グリッドに及ぶことになります。しかし、今回の分析は「ラボ内での差異」のみを重視しているため、この点は結果に影響しません。
- SVGmaxxing(SVG最適化)の影響: あるラボが SVG 生成全体、あるいは動物と車両の組み合わせのような特定のサブセットに対して最適化を行った場合、そのラボはすべてのセルでスコアが上昇し、単に「上手な」ラボと同じように見えてしまいます。Google/DeepMind などの一部のラボはこの手法を公言しています。しかし、今回の実験ではそのような傾向を検出することはできません。
- 予算は限られていました。実験全体で API クレジット約 80 ドルしか使えなかったのです。これにより、1 セルあたり 3 サンプルずつ、審査員は 1 名、モデルは 7 つに制限されました。また、「飛行機」対「航空機」といったケースのように、プロンプトやパイプラインの微調整を何度も繰り返すことも防がれました。
結論
HN の批判派の方々、ごめんなさい。AI ラボが意図的にペリカン画像を生成しているという証拠はほとんどありません。少なくとも、明らかに目立つようなやり方では行っていないようです。
ペリカンの描画が他の動物より優れているわけではありません。自転車の描画も他の乗り物より優れていません。そして、どのラボも「ペリカンと自転車」の組み合わせを、それぞれのモデルが予測する以上にうまく描けているわけではありません。GLM-5.2 が最も近い結果を出しました。この特定のセルでのスコア上昇幅は最大で、最初の「ペリカンが自転車に乗っている」サンプルにも目を引くものがありました。しかし、その効果は小さく統計的に有意ではないため、過度に重視する必要はないでしょう。
もう一つ目立つのは、シーンの構成における方向性です。21 枚の「ペリカンと自転車」の画像すべてが右向きになっています。これはグリッド内での唯一のケースで、すべての画像が一致しています。しかし、それほど不思議なことでもありません。実験全体を通して右向きの描画が標準的だったのです。他の組み合わせでも 3 つが 90% 以上の一致率を記録しており、48 種類もの組み合わせがある中で、1 つが 21 枚中 21 枚すべてで一致したとしても驚きません。
より可能性が高いのは、Google や DeepMind が行っているような「SVGmaxxing」です。他のラボも同様のことを静かに行っているかもしれません。残念ながら、今回の実験では誰が行っているのかを特定することはできません。しかし、少なくとも今夜は安心してください。AI ラボが Simon Willison を騙すためにペリカンと自転車の画像をテラバイト単位で生成しているわけではないのですから。
ご自身でもデータを参照したい場合は、完全なパイプラインを リポジトリ で公開しています。
脚注
- 人気のあるベンチマークで高いスコアを獲得するために AI モデルを最適化する実践のこと。↩︎
引用
BibTeX 形式の引用情報:
@online{castillo2026,
author = {Castillo, Dylan},
title = {Are {AI} Labs Pelicanmaxxing?},
date = {2026-07-18},
url = {https://dylancastillo.co/posts/pelicanmaxxing.html},
langid = {en}
}
本稿を引用する際は、以下のように記載してください。
Castillo, Dylan. 2026. "Are AI Labs Pelicanmaxxing?" July 18. https://dylancastillo.co/posts/pelicanmaxxing.html.
原文を表示
For the past few years, Simon Willison has tested every major LLM release with the same prompt: “Generate an SVG of a pelican riding a bicycle”.
What began as a tongue-in-cheek benchmark has become one of the most famous informal benchmarks in AI. Simon’s pelican-on-a-bicycle results are often among the most upvoted comments on Hacker News threads announcing new releases from AI labs.
The benchmark is now famous enough that there’s plenty of discussion about its usefulness and about whether AI labs might be benchmaxxing1 on it. When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn’t it be tempting to *pelicanmaxx* your model just a bit?
I wanted to find out, so I put together a small experiment. I generated 1,008 SVGs across seven frontier models, scored them with an LLM judge, and used Claude Fable 5 for the analysis.
This article presents the results. All the code is available on Github.
How I tested it
I built a grid of 8 animals × 6 vehicles = 48 prompts, where the famous prompt is one cell:
- Animals: pelican, flamingo, heron, otter, raccoon, antelope, whale, cat
- Vehicles: bicycle, unicycle, skateboard, scooter, plane, boat
Every prompt uses almost identical phrasing to Simon’s, only switching the animal and vehicle. The animal and vehicle selection wasn’t done in a very rigorous manner, but I tried to vary both similarity to the original prompt and difficulty. Flamingo and heron are quite similar to pelicans; cat, raccoon, and otter are easy cases; antelope is hard; and whale is as different as you can get.
I tested seven models through OpenRouter: *GPT-5.6 Terra*, *Claude Sonnet 5*, *Gemini 3.5 Flash*, *Grok 4.5*, *Qwen3.7-Max*, *GLM-5.2*, and *DeepSeek V4 Pro*. I generated 3 samples per prompt, at temperature 1.0, requesting the same reasoning effort from every model. That resulted in 1,008 SVGs.
Then I ran each image through a three-stage pipeline:
- Rendering: Each SVG is rendered to PNG. If a model returns no SVG or one that fails to render, I regenerate until it produces a valid one, and record the number of attempts. There were only 11 retries across the 1,008 generations.
- Judging: GPT-5.6 Luna scores each image with 1-5 ratings for the animal, the vehicle, and the coherence of the action. When I rank animals or vehicles below, I use the matching rating on its own. When I need one number per image, I use the average of the three, which I call the judge score.
- Feature extraction: For a more detailed analysis, I also passed each rendered image to Gemini 3.1 Flash-Lite, which recorded the animal and vehicle it recognized, which way the subject faces, and an open-ended list of scene elements.
My hypothesis is that if a lab trained on the benchmark, it should show up in some combination of the pelican row scoring above what the animal deserves, the bicycle column scoring above what the vehicle deserves, or the specific pelican-bicycle cell beating both.
Evidence #1: The pelicans on bicycles don’t look any better
Before any scoring, the simplest test is to look at the images yourself. Pick a lab to see everything it drew, with the judge’s score under each image (click to open full size):
LabSample
pelican
bicycle
failed
A4 · V5 · C4
unicycle
failed
A5 · V4 · C5
skateboard
failed
A5 · V4 · C5
scooter
failed
A4 · V4 · C3
plane
failed
A4 · V5 · C5
boat
failed
A5 · V5 · C5
flamingo
bicycle
failed
A5 · V4 · C3
unicycle
failed
A4 · V4 · C4
skateboard
failed
A4 · V5 · C3
scooter
failed
A5 · V5 · C5
plane
failed
A4 · V5 · C3
boat
failed
A4 · V5 · C4
heron
bicycle
failed
A5 · V5 · C4
unicycle
failed
A4 · V4 · C5
skateboard
failed
A4 · V5 · C5
scooter
failed
A5 · V5 · C5
plane
failed
A5 · V1 · C1
boat
failed
A5 · V5 · C5
otter
bicycle
failed
A5 · V4 · C2
unicycle
failed
A5 · V3 · C3
skateboard
failed
A5 · V5 · C5
scooter
failed
A5 · V4 · C5
plane
failed
A3 · V5 · C5
boat
failed
A4 · V5 · C4
raccoon
bicycle
failed
A5 · V3 · C4
unicycle
failed
A5 · V3 · C3
skateboard
failed
A5 · V5 · C5
scooter
failed
A5 · V5 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
antelope
bicycle
failed
A5 · V5 · C3
unicycle
failed
A4 · V4 · C4
skateboard
failed
A4 · V5 · C4
scooter
failed
A4 · V5 · C3
plane
failed
A5 · V5 · C4
boat
failed
A5 · V5 · C5
whale
bicycle
failed
A5 · V3 · C3
unicycle
failed
A5 · V3 · C3
skateboard
failed
A5 · V4 · C5
scooter
failed
A5 · V5 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
cat
bicycle
failed
A5 · V3 · C3
unicycle
failed
A5 · V4 · C3
skateboard
failed
A5 · V5 · C5
scooter
failed
A5 · V4 · C4
plane
failed
A5 · V5 · C5
boat
failed
A5 · V5 · C5
I looked through the images myself before running the analysis below. Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid. Maybe in GLM-5.2’s first sample it felt slightly better than the rest, but that batch also produced a pretty cool heron on a skateboard, so I cannot say for sure. Otherwise they look like the rest of what each model draws, and the labs that draw good pelicans on bicycles also do a good job drawing other animal-vehicle combinations.
But this test is hard to replicate, and everyone will have a different opinion. So I wanted something more quantitative, which is why I opted for the method detailed above.
Evidence #2: Labs are not better at drawing pelicans
Here’s the mean animal rating per animal, pooled across all models:
The pelican is 6th of 8, behind cat, whale, raccoon, heron, and antelope. If AI labs were training on the benchmark, you’d expect pelicans at the top. Instead they’re in the bottom half. All seven labs draw cats, whales, and raccoons better than pelicans.
Of course, a pelican may simply be harder to draw than a cat. A lab could train on pelicans and still not push them past the easy animals, so this ranking alone can’t rule that out. I’ll adjust for difficulty in Evidence #4.
Evidence #3: Labs are not better at drawing bicycles
Bicycles fare even worse. They sit second from last, in a near-tie with planes, which come in last:
If labs were training on the benchmark, you’d expect bicycles near the top of this ranking. They’re not. However, the same caveat applies here. A bicycle is harder to draw than a skateboard: it needs two matching wheels, a frame that reaches both axles, handlebars, a seat, and pedals. The judge flags a missing or disconnected one of those on 2/3 of the bicycle images. You can train on bicycle images and still not do a great job relative to simpler vehicles.
One note on the plane, though: I should’ve picked “airplane” instead of “plane” because models often read it geometrically. They drew the animal standing on a flat surface instead of flying an aircraft. The plane is the only vehicle where the feature extractor sometimes found no vehicle at all (25 of 168 images, against zero for the other five), and 20% of plane images scored a 1 or 2 on the vehicle rating, against 5% for bicycles and none at all for boats, scooters, or skateboards.
Evidence #4: Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty
Put the two together and the “pelican on a bicycle” ends up near the bottom of the ranking, at #42 of 48:
But again, some combinations might be just harder to draw than others.
To account for that, I fit a fixed-effects regression on all 1,008 images: score ~ lab + animal × vehicle, plus per-lab interaction terms for pelican, bicycle, and the pelican-bicycle cell, with robust standard errors. The animal × vehicle terms absorb the inherent difficulty of all 48 combinations. The interactions measure each lab’s benchmark-specific boost relative to the average lab, with confidence intervals.
The results:
- Every per-lab pelican effect (the lab’s boost on pelicans across all six vehicles) lands between -0.11 and +0.14 judge points, and none comes close to significance (smallest p = 0.25).
- The per-lab bicycle effects (the lab’s boost on bicycles across all eight animals) run from Grok 4.5 at -0.18 (p=0.11) to Gemini 3.5 Flash at +0.27 (p=0.022). Only Gemini clears p < 0.05, and the seven point in both directions.
- No pelican-bicycle cell effect (the extra boost on the specific combination, on top of the lab’s pelican and bicycle effects) clears p < 0.05. The largest positive is GLM-5.2 at +0.35 (p=0.12), which is the one I mentioned earlier. It’s the closest thing to a signal in this experiment, but still within chance.
Here are the full per-lab estimates. A pelicanmaxxing lab would show dots to the right of the zero line across its whole row:
Every pelican interval and every cell interval contains zero. Exactly one doesn’t: *Gemini 3.5 Flash* in the bicycle column. But with 21 tests at p < 0.05, chance alone predicts about one false positive (21 × 0.05 ≈ 1.05), and one is exactly what came up. It also doesn’t survive a multiple-comparisons correction: the Bonferroni threshold across the 21 tests is 0.05/21 ≈ 0.002, and its p-value is 0.022. The full table of estimates and p-values is in the repo.
But these intervals are wide, about ±0.6 judge points on average. Any boost smaller than that won’t be captured by this test.
Evidence #5: The pelican-bicycle scenes don’t look memorized
Some have suggested that the pelican on a bicycle looks like a memorized composition, pointing to recurring patterns such as the pelican always facing right, or recurring elements like a sun or a scarf. So I wanted to know if this was true.
Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.
However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest:
vehicle
left
right
ambiguous
scooter
9%
83%
8%
bicycle
11%
81%
8%
skateboard
19%
60%
21%
plane
21%
58%
21%
unicycle
22%
45%
33%
boat
33%
35%
32%
Pelicans are also among the animals that tend to face right:
animal
left
right
ambiguous
antelope
21%
78%
1%
pelican
22%
78%
0%
heron
22%
77%
1%
whale
34%
65%
1%
flamingo
36%
64%
0%
otter
8%
45%
47%
cat
8%
40%
52%
raccoon
3%
36%
61%
It’s hard to draw a pelican or a bicycle facing the viewer, so models almost always draw them from the side, facing left or right. That’s why so few of their images are ambiguous. Other combinations also come close to unanimous: antelope on a scooter and pelican on a scooter land at 20 of 21, and heron on a bicycle at 19 of 21. So 21 out of 21 doesn’t seem like an outlier.
Scene elements: I let the extractor name any element it saw in the image. These are the counts:
A memorized scene would show up as the same set of elements recurring picture after picture. I went looking for that, and found some combinations do tend to produce the same elements every time. Every single flamingo on a boat has a sun in it. Otters on planes wear scarves 38% of the time. Cats on bicycles get a basket 38% of the time.
The pelican on a bicycle doesn’t seem to have anything particularly different about it. It just has some elements that appear more frequently, like every other animal-vehicle combination.
Limitations
- Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run. If a model just can’t judge a drawing reliably, none of the numbers above mean much. The judge is also from the same family as one of the contestants, GPT-5.6 Terra. However, every lab draws all 48 combinations, so a judge that happens to like one lab’s style lifts that lab’s whole grid at once. But that doesn’t change the results because this analysis only cares about the within-lab differences.
- SVGmaxxing. A lab that optimized SVG generation as a whole (or a subset such as animals on vehicles) rises on every cell at once and looks identical to a lab that’s just good. Some labs, such as Google/DeepMind, openly do this. This experiment can’t detect that.
- Limited budget. The whole experiment ran on roughly $80 of API credits. That capped it at 3 samples per cell, a single judge, and 7 models. This also prevented me from iterating too much on the prompts and pipeline, as with the “plane” vs. “airplane” case.
Conclusion
Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing. Or at least they’re not doing it in a plainly obvious manner.
Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.
The other thing that stands out is direction in the scene composition. All 21 pelican-bicycle images face right, the only combination in the grid where every image agrees. But it doesn’t seem that strange. Facing right is the norm across the experiment. Three other combinations land at 90% or above, and with 48 of them, I’m not surprised one reached 21 out of 21.
The more plausible story is SVGmaxxing like Google/DeepMind does. Other labs might be doing it more quietly. Sadly, this experiment can’t say who’s doing it. But at least you can sleep tonight knowing that AI labs are not producing terabytes of pelicans on bicycles just to trick Simon Willison.
If you want to look at the data yourself, the full pipeline is in the repo.
Footnotes
- the practice of optimizing AI models to achieve high scores on popular benchmarks.↩︎
Citation
BibTeX citation:
@online{castillo2026,
author = {Castillo, Dylan},
title = {Are {AI} Labs Pelicanmaxxing?},
date = {2026-07-18},
url = {https://dylancastillo.co/posts/pelicanmaxxing.html},
langid = {en}
}
For attribution, please cite this work as:
Castillo, Dylan. 2026. “Are AI Labs Pelicanmaxxing?” July
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み