Qwen3.5-Omniが音声指示と映像からコードを書く方法を誰にも教わらずに習得
本文の状態
日本語全文を表示中
詳細モードで約1分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
The Decoder
アリババが音声・映像・画像・テキストを処理する多モーダルAIモデル「Qwen3.5-Omni」を発表した。同モデルは音声タスクでGemini 3.1 Proを上回り、訓練なしに音声指示と映像入力からコードを生成する能力を獲得した。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。

Alibabaは、テキスト、画像、音声、動画を処理するオムニモーダルAIモデル「Qwen3.5-Omni」をリリースしました。同社は、このモデルが音声タスクにおいてGemini 3.1 Proを上回ったと主張しており、開発過程で予期せぬ能力を獲得しました。それは、音声指示と動画入力からコードを記述することです。
記事「Qwen3.5-Omni learned to write code from spoken instructions and video without anyone training it to」は、The Decoderで最初に公開されました。
原文を表示

Alibaba has released Qwen3.5-Omni, an omnimodal AI model that processes text, images, audio, and video. It claims to beat Gemini 3.1 Pro on audio tasks and picked up an unexpected trick along the way: writing code from spoken instructions and video input.
The article Qwen3.5-Omni learned to write code from spoken instructions and video without anyone training it to appeared first on The Decoder.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み