Together AI、Moonshot AI と戦略的提携し Kimi モデルをネイティブ提供へ
Together AI は Moonshot AI と戦略的提携を結び、2.8兆パラメータのオープンモデル「Kimi K3」および今後の全オープンモデルを同社のプラットフォーム上でネイティブに提供すると発表した。
AI深層分析を開く2026年7月30日 05:49
AI深層分析
キーポイント
Moonshot AI との戦略的提携発表
Together AI は Moonshot AI と提携し、Kimi K3 を皮切りに Moonshot がリリースするすべてのオープンウェイトモデルを Together AI のプラットフォーム上でネイティブに提供すると発表した。
Kimi K3 モデルの技術的特徴
Kimi K3 は 2.8 兆パラメータのスプース MoE アーキテクチャを持ち、KDA と AttnRes という新コンポーネントにより長文脈での推論速度と効率が大幅に向上している。
開発者向けインフラと機能
Together AI の US ホスト型インフラ上で、データ保持ゼロの条件で Kimi K3 にアクセスでき、独自データでのポストトレーニングも可能となる。
プロダクション対応プラットフォーム
Serverless や Provisioned Throughput などの製品を通じて、SLA とオートスケーリングを備えた環境で Kimi モデルを実運用に適用できる。
プロダクション向けインフラとSLAの提供
Provisioned Throughputにより、トークンベースの価格設定で99%の稼働率を保証する予約推論容量を利用可能になる。Dedicated Model Inferenceでは高速な展開と継続的な研究革新が製品に組み込まれる。
重要な引用
Kimi K3, Moonshot AI's 2.8-trillion-parameter open model is live on Together AI
Together AI becomes a launch platform for Moonshot's model releases
Moonshot says combine to deliver roughly 2.5x better scaling efficiency compared to Kimi K2
It's the reserved-capacity guarantee developers already expect from closed-model providers, now available for open weights.
編集コメントを表示
編集コメント
2.8 兆パラメータ規模のオープンモデルが実用化されるのは、業界にとって大きな転換点となる。Together AI のインフラ上でデータ保持ゼロを担保しつつ利用可能になる点は、企業利用における信頼性を高める重要な要素である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Moonshot AI が開発したパラメータ数 2.8 兆のオープンモデル「Kimi K3」が、Together AI で利用可能になりました。今後は Moonshot から公開されるすべてのオープンモデルも同様に提供していく方針です。
本日、私たちは大規模な MoE(Mixture of Experts)アーキテクチャでオープンソースの最前線を切り拓く主要なモデル開発ラボの一つである Moonshot AI と、戦略的パートナーシップを締結したことを発表します。この提携により、Together AI は Moonshot のモデルリリースプラットフォームとして機能します。まずは Kimi K3 から始まり、今後 Moonshot が公開するすべてのオープンウェイトモデルが対象となります。
オープンモデルを活用してアプリケーションを開発している開発者にとって、これは Together AI が提供する米国ホストのインフラストラクチャ、価格体系、およびツールを通じて、Moonshot の最新リリースを発売当日から利用可能になることを意味します。同時に、データの保持ゼロとコンプライアンスに準拠したアクセスも保証されます。さらに、開発者は自社のデータでこれらのモデルを追加学習(post-train)することで、自社アプリに必要な品質とパフォーマンスを実現できます。
最前線の性能とオープンウェイト
Kimi K3 はこれまでに発表された中で最も大規模なオープンモデルです。ネイティブのビジョン機能を備え、100 万トークンのコンテキストウィンドウを有する、パラメータ数 2.8 兆のスプース MoE モデルです。Moonshot はこれを、2 つの新規アーキテクチャコンポーネントを中心に構築しました。
Kimi Delta Attention (KDA) は、シーケンス長にわたる情報の流れを変革し、長いコンテキスト長において大幅なデコード速度の向上を実現します。
Attention Residuals (AttnRes) は、モデル深さ全体での表現の取得方法を改善し、最小限の追加計算コストで意味のあるトレーニング効率をもたらします。
これらの上に層を成すように配置されたのが、最適化手法の改良群です。ヘッドごとの Muon や、量子化に基づくエキスパート負荷分散などが含まれます。Moonshot によると、これらの組み合わせにより、Kimi K2 と比較して約 2.5 倍のスケーリング効率が達成されます。
その結果、長期にわたるコーディング、エージェントワークフロー、ゲーム開発、知識集約型タスクに対応したモデルが誕生しました。ベンチマーク結果は、現在市場にある主要な独自システムと直接競合するレベルにあります。
本番環境向け推論プラットフォーム
Kimi モデルは、Together AI の推論製品群(Serverless、Provisioned Throughput、Dedicated Inference)を通じて利用可能です。開発者は、これらのモデルを試すことから、SLA を保証した製品や自動スケーリング機能を活用して、本番環境のユースケースへの適用までをスムーズに進めることができます。
Together AI の「プロビジョンド・スループット(Provisioned Throughput)」を利用すれば、開発チームは先端的なオープンモデルに対して予約済み推論リソースを確保できます。トークンベースの課金体系を採用しつつ、99% の稼働率を保証する SLA も提供されます。これは従来クローズドモデルのプロバイダーから当然とされていた「容量保証」が、オープンウェイトモデルでも利用可能になったことを意味します。GPU 時間の計算に悩まされる必要はなく、予測可能な価格で確実なスループットが得られます。
より高い制御性と、本番環境で即座に使えるプラットフォームの利点を両立させたいチームには、「専用モデル推論(Dedicated Model Inference)」の利用もおすすめです。迅速なデプロイが可能で、トークンコストの最適化や、製品へ継続的に取り込まれる研究開発の最新成果など、多くのメリットがあります。
Post Training
開発者は Kimi モデルをポストトレーニングすることで、特定のユースケースに合わせた高品質なモデルを提供できます。カスタムトレーニングでは、フルウェイト学習や LoRA による強化学習、高度な教師あり微調整(SFT)などに対応可能です。Python SDK や細粒度のプリミティブを通じてトレーニング設定を行い、専用リソース上で複数の LoRA 実験を並列実行することもできます。このようにして、実験と本番環境が直接結びつき、開発サイクルが加速します。
評価用のチェックポイントが完成すれば、そのまま推論環境へネイティブにデプロイ可能です。トレーニングとサービス提供の間の手動引き継ぎは不要で、モデルを再構築する必要もありません。
What this partnership unlocks for developers
Together AI は Moonshot AI と戦略的パートナーシップを結び、Kimi モデルをネイティブに提供することを発表しました。
Day zero(リリース直後)から Kimi K3 が Together AI で利用可能になります。Moonshot 社が検証した最高品質とパフォーマンスを実現し、今後の Moonshot 製モデルも発売と同時に Together AI で提供されます。
実証済みのスケールと本番環境対応インフラを備えています。Together AI の研究最適化された推論スタックは、K3 のような大規模スパース MoE モデル向けにチューニングされており、開発者はアーキテクチャが約束する性能をそのまま享受できます。一般的なデプロイではなく、Cursor や Y Combinator、Decagon といった企業が既に本番トラフィックで利用している実績があります。これにより、コード生成やエージェント処理などの大規模ワークロードをスケールして展開可能です。
1 つの統合で Moonshot の全ラインナップが利用可能になります。Moonshot が新モデルを発表するたびに、それらは同じ Together AI モデルライブラリに追加され、すでに使用している API をそのまま呼び出すだけで済みます。新しい SDK の導入や別アカウントの作成、Kimi 新モデルのリリースごとにアプリを再構築する必要はありません。
Day zero からポストトレーニングも利用可能です。Together の顧客は Kimi モデルをベースモデルとしてファインチューニングに活用でき、MIT ライセンスで要求される出典表示の削除も柔軟に行えます。
オープンウェイトにより、制御権はユーザーが選択できます。Kimi K3 はオープンウェイトであるため、Together 上でサーバーレス実行して迅速な反復開発が可能ですが、利用規模が大きくなればプロビジョニングスループットや専用モデル推論へ移行も可能です。特定のベンダーのロードマップに縛られることなく、柔軟に対応できます。
議論に参加しよう
Moonshot AI と Together AI のチームメンバーから直接話を聞きたい方へ。Moonshot AI の Feihu Tang 氏と、Together AI の Jue Wang 氏、Zain Hasan 氏が登壇し、Kimi K3 の技術的な背景や、Together 上で実際に運用する方法、今後の展開について議論します。ご質問もお待ちしています。
原文を表示
Kimi K3, Moonshot AI's 2.8-trillion-parameter open model is live on Together AI, with a commitment to land all future open Moonshot models.
Today we're announcing a strategic partnership with Moonshot AI, one of the leading model labs pushing the open source frontier with large-scale MoE architectures. Under this partnership, Together AI becomes a launch platform for Moonshot's model releases, starting with Kimi K3 and extending to every open weights model Moonshot ships going forward.
For developers building on open models, this means day zero access to Moonshot's frontier releases through Together AI’s US-hosted infrastructure, pricing model, and tooling, while ensuring zero data retention and compliant access to models and data. Developers can also post train these models with their own data to deliver the quality and performance for their app.
Frontier performance, open weights
Kimi K3 is the largest open model released to date: a 2.8T parameter sparse Mixture-of-Experts model with native vision support and a 1M token context window. Moonshot built it around two new architectural components:
- Kimi Delta Attention (KDA), which changes how information flows across sequence length and delivers significantly faster decoding at long context lengths.
- Attention Residuals (AttnRes), which improves how representations are retrieved across model depth, adding meaningful training efficiency at minimal extra compute cost.
Layered on top are a set of optimizer refinements, including per-head Muon and quantile-based expert load balancing, that Moonshot says combine to deliver roughly 2.5x better scaling efficiency compared to Kimi K2. The result is a model built for long-horizon coding, agentic workflows, game development, and knowledge-intensive tasks, with benchmark results that put it in direct competition with the leading proprietary systems on the market today.
Production Inference Platform
Kimi models are available across Together AI Inference products – including Serverless, Provisioned Throughput and Dedicated Inference. Developers can go from trying these models to applying them to their production use cases with SLA-backed products and autoscaling.
With Provisioned Throughput, teams can use reserved inference capacity for frontier open models with token-based pricing and a 99% uptime SLA. It’s the reserved-capacity guarantee developers already expect from closed-model providers, now available for open weights. No GPU-hour math, just guaranteed throughput at a predictable price. For teams looking for more control with all the benefits of a production-ready platform, they can use Dedicated Model Inference with fast deployment, better token economics and continuous research innovations shipped into the product.
Post Training
Developers can also post train Kimi models to deliver better quality for their target use case. With custom training, including full-weight and LoRA reinforcement learning as well as advanced supervised fine-tuning, developers can configure training through the Python SDK and granular primitives, and run multiple LoRA experiments concurrently on dedicated capacity. Custom training connects experimentation directly to production. When a checkpoint is ready to evaluate, it can be deployed natively to inference, with no separate handoff between training and serving and no rebuilding.
What this partnership unlocks for developers
- Day zero availability: Kimi K3 is available on Together AI with the highest model quality and performance, validated by Moonshot. Future Moonshot releases will also ship on Together at launch.
- Proven scale with production-ready infrastructure: Together AI's research-optimized inference stack is tuned for large sparse MoE models like K3, so developers get the performance the architecture promises rather than a generic deployment. Together’s inference stack already serves production traffic for companies like Cursor, Y Combinator and Decagon, deploying coding and agentic workloads at scale.
- One integration, the full Moonshot lineup: As Moonshot ships new models, they'll land in the same Together AI Models library, behind the same API you're already calling. No new SDKs, no separate account, no re-plumbing your app each time a new Kimi model drops.
- Day zero post training: Together customers can use Kimi models as their base model for fine-tuning, with the flexibility to remove the attribution required in the MIT license.
- Open weights, your choice of control: Because Kimi K3 is open-weight, you can run it serverless on Together for speed of iteration, or move to Provisioned Throughput or Dedicated Model Inference as your usage scales, without being locked into a single provider's roadmap.
Join the conversation
Want to hear directly from the team behind Moonshot AI and Together AI? Join Feihu Tang from the Moonshot AI team and Jue Wang and Zain Hasan from Together AI to talk through the technical decisions behind Kimi K3, how to actually run it on Together, and what's coming next. Bring your questions.
AI算出
主要ニュースainew評価高い
AI モデルのアーキテクチャ詳細(KDA, AttnRes)やベンチマーク結果、推論プラットフォームの新機能など、技術的な深みのある新規情報を含まれており、単なる提携発表以上の価値がある。ただし、日本企業との直接的な関わりや日本固有の情報が欠けているため、日本の関連性は低めとなる。
6つの評価軸を見る
- AI関連度
- 100
- 情報源の信頼性
- 100
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み