JetBrains の Mellum 2(49 分読み)
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
JetBrains が開発ツール「Mellum」のバージョン 2 を公開し、詳細な機能解説を 49 分間の読了量で提供している。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
**
要約:Mellum 2 を発表します。これは、1 トークンあたり 25 億パラメータが活性化するオープンウェイトの 120 億パラメータ Mixture-of-Experts (MoE) 言語モデルです。Mellum 2 はソフトウェアエンジニアリングに特化した汎用言語モデルであり、コード生成と編集、デバッグ、多段階推論、ツール使用および関数呼び出し、エージェント型コーディング、対話型プログラミング支援を網羅しています。これは完了指向の 40 億パラメータ密度モデルである Mellum の後継モデルです。アーキテクチャは Mixture-of-Experts (64 エキスパート中 8 が活性) を基盤とし、Grouped-Query Attention (4 つの KV ヘッド付き)、4 レイヤーに 3 レイヤー分の Sliding Window Attention、および補助的な事前学習目的と推測デコーディング用の内蔵ドラフトモデルの両方として機能する単一の Multi-Token Prediction ヘッドを組み合わせました。各選択は、コモディティ GPU における推論効率を設計制約としたアブレーション検証によって裏付けられています。事前学習は、多様なウェブデータから厳選されたコードおよび数学コンテンツへと段階的に混合比率をシフトさせる 3 フェーズのカリキュラムを通じて、約 10.6 トリリオントークンにわたって行われます。これは Muon オプティマイザを用いて FP8 ハイブリッド精度で最適化され、線形減衰でゼロに至る Warmup-Hold-Decay スケジュールで学習されます。事前学習済みベースモデルは、レイヤー選択型の YaRN を用いて 128K コンテキストウィンドウに拡張され、その後 2 つの段階(教師あり微調整に続く RLVR)でポストトレーニングが行われます。その結果、直接回答する Instruct モデルと、最終回答の前に明示的な推論トレースを出力する Thinking モデルという 2 つのリリース版が得られます。コード生成、数学・推論、ツール使用、知識、安全性の各ベンチマークにおいて、Mellum 2 は 40 億〜140 億パラメータ範囲のオープンウェイトベースラインと競合する性能を示しつつ、25 億パラメータ密度モデルに相当するトークンあたりの計算量で動作します。私たちは、アーキテクチャ決定、データパイプライン、トレーニングレシピに関する本報告書とともに、ベース、インストラクション、思考の各チェックポイントを Apache 2.0 ライセンスの下で公開します。
**
主題:
コンピュータ言語 (cs.CL)
引用形式:
arXiv:2605.31268 [cs.CL]
(またはこのバージョンについては
arXiv:2605.31268v1 [cs.CL])
https://doi.org/10.48550/arXiv.2605.31268
DataCite 経由の arXiv 発行 DOI
## 提出履歴
From: Nikiita Pavlichenko [メールを表示]
[v1]**
2026 年 5 月 29 日 (金) 13:01:11 UTC (1,508 KB)
原文を表示
Abstract:We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-purpose language model specialized in software engineering, spanning code generation and editing, debugging, multi-step reasoning, tool use and function calling, agentic coding, and conversational programming assistance, and it is the successor to the completion-focused 4B dense Mellum model. The architecture builds on the Mixture-of-Experts (64 experts, 8 active) and combines Grouped-Query Attention with 4 KV heads, Sliding Window Attention on three of every four layers, and a single Multi-Token Prediction head that doubles as both an auxiliary pre-training objective and a built-in draft model for speculative decoding; each choice was validated by ablation with inference efficiency on commodity GPUs as a design constraint. Pre-training spans approximately 10.6 trillion tokens through a three-phase curriculum that progressively shifts the mixture from diverse web data toward curated code and mathematical content, optimized with Muon under FP8 hybrid precision and a Warmup-Hold-Decay schedule with linear decay to zero. The pre-trained base is extended to a 128K context window via a layer-selective YaRN and then post-trained in two stages (supervised fine-tuning followed by RLVR), yielding two released variants: an Instruct model that answers directly and a Thinking model that emits an explicit reasoning trace before its final answer. Across code generation, math and reasoning, tool use, knowledge, and safety benchmarks, Mellum 2 is competitive with open-weight baselines in the 4B-14B range while running at the per-token compute of a 2.5B dense model. We release the base, instruct, and thinking checkpoints, together with this report on the architecture decisions, data pipeline, and training recipe behind them, under the Apache 2.0 license.
| Subjects: | Computation and Language (cs.CL) |
|---|---|
| Cite as: | arXiv:2605.31268 [cs.CL] |
| (or arXiv:2605.31268v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2605.31268 arXiv-issued DOI via DataCite |
Submission history
From: Nikiita Pavlichenko [view email] [v1]
Fri, 29 May 2026 13:01:11 UTC (1,508 KB)
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み