Qwen3.6-35B-A3BがClaude Opus 4.7より優れたペリカン画像を生成
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
6媒体で確認
Simon Willison Blog · 宝玉的分享 · InfoQ · AWS Machine Learning Blog · TechCrunch AI · The Decoder
各社の報じ方を比較 ↓著者が公開した自転車に乗るペリカンのベンチマークテストで、AlibabaのQwen3.6-35B-A3BがAnthropicのClaude Opus 4.7より優れた画像を生成したことを報告している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
もし私のpelican riding a bicycle benchmarkを、モデルをテストする堅牢な手法として真剣に受け止めてくださっている方がいれば、今朝の2つの大規模モデルリリースからのペリカン画像をお届けします。Alibaba製のQwen3.6-35B-A3Bと、Anthropic製のClaude Opus 4.7です。
こちらはQwen 3.6のペリカンです。Unslothによるthis 20.9GB Qwen3.6-35B-A3B-UD-Q4_K_S.gguf量子化モデル (quantized model) を使用し、MacBook Pro M5上でLM Studio(およびllm-lmstudioプラグイン)経由で実行して生成しました。-transcript here:

そしてこちらはAnthropicのbrand new Claude Opus 4.7から得たものです。(transcript):

こちらはQwen 3.6に軍配を上げます。Opusは自転車フレームの描写を台無しにしてしまいました!
thinking_level: max (thinking_level) を渡してOpusを2回目のテストで試しましたが、あまり改善しませんでした(transcript):

Qwenが不正をしているとは思えない
多くの人が、ラボが私のくだらないベンチマークのために学習していると確信しています。私はそうは思いませんが、正直この結果は少し疑念を抱かせるものでした。そこで秘密のバックアップテストの一つを投入します - 「Generate an SVG of a flamingo riding a unicycle」というプロンプトで、Qwen3.6-35B-A3BとOpus 4.7から得た結果がこちらです:
Qwen3.6-35B-A3B
image
Opus 4.7
image
私はこの結果もQwenに軍配を上げます。優れたSVGコメント (SVG comment) のおかげもあってのことです。
ここから何が学べるのか
このpelican benchmarkは常にジョークとして意図されたものであり、主にこれらのモデルを比較するという作業がいかに難解で不合理であるかを示す声明です。
そのジョークの奇妙な点は、大部分において、生成されたペリカンの品質とモデル (models) の全体的な有用性の間に直接的な相関関係があったことです。それらの2024年10月最初のペリカンはゴミ同然でした。より最近の投稿は全体的にずっとずっと良く、ペリカンが自転車に乗っているイラストを急務で必要とするような場合であれば、Gemini 3.1 Proは実際にどこかで使えるイラストを生み出すに至っています。
今日の時点で、その有用性への緩やかな結びつきでさえ断ち切られています。Qwenに対しては多大な敬意を抱いていますが、彼らの最新モデルの21GB量子化版 (quantized version) がAnthropicの最新の独自リリース (proprietary release) よりも強力であるか、あるいは有用であると確信できるわけではありません。
ただし、あなたが求めているのがペリカンが自転車に乗っているSVGイラスト (SVG illustration) であるならば、現時点ではラップトップ上で動作するQwen3.6-35B-A3Bの方が、Opus 4.7よりも確実な選択肢と言えます!
タグ: ai, generative-ai, local-llms, llms, anthropic, claude, qwen, pelican-riding-a-bicycle, llm-release, lm-studio
原文を表示
For anyone who has been taking my pelican riding a bicycle benchmark seriously as a robust way to test models, here are pelicans from this morning's two big model releases - Qwen3.6-35B-A3B from Alibaba and Claude Opus 4.7 from Anthropic.
Here's the Qwen 3.6 pelican, generated using this 20.9GB Qwen3.6-35B-A3B-UD-Q4_K_S.gguf quantized model by Unsloth, running on my MacBook Pro M5 via LM Studio (and the llm-lmstudio plugin) - transcript here:

And here's one I got from Anthropic's brand new Claude Opus 4.7 (transcript):

I'm giving this one to Qwen 3.6. Opus managed to mess up the bicycle frame!
I tried Opus a second time passing thinking_level: max. It didn't do much better (transcript):

I don't think Qwen are cheating
A lot of people are convinced that the labs train for my stupid benchmark. I don't think they do, but honestly this result did give me a little glint of suspicion. So I'm burning one of my secret backup tests - here's what I got from Qwen3.6-35B-A3B and Opus 4.7 for "Generate an SVG of a flamingo riding a unicycle":


I'm giving this one to Qwen too, partly for the excellent `` SVG comment.
What can we learn from this?
The pelican benchmark has always been meant as a joke - it's mainly a statement on how obtuse and absurd the task of comparing these models is.
The weird thing about that joke is that, for the most part, there has been a direct correlation between the quality of the pelicans produced and the general usefulness of the models. Those first pelicans from October 2024 were junk. The more recent entries have generally been much, much better - to the point that Gemini 3.1 Pro produces illustrations you could actually use somewhere, provided you had a pressing need to illustrate a pelican riding a bicycle.
Today, even that loose connection to utility has been broken. I have enormous respect for Qwen, but I very much doubt that a 21GB quantized version of their latest model is more powerful or useful than Anthropic's latest proprietary release.
If the thing you need is an SVG illustration of a pelican riding a bicycle though, right now Qwen3.6-35B-A3B running on a laptop is a better bet than Opus 4.7!
Tags: ai, generative-ai, local-llms, llms, anthropic, claude, qwen, pelican-riding-a-bicycle, llm-release, lm-studio
同じ出来事を6媒体で確認
同じ出来事を扱う別媒体の記事です。見出しと公開時刻を比較できます。
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み