Writer、新モデル Palmyra X6 で AI エージェントコストを 52% 削減
本文の状態
日本語全文を表示中
詳細モードで約19分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
Writer は中国企業 Z.ai のオープンソースモデル GLM-5.2 を基盤に自社のインフラで再訓練した「Palmyra X6」を発表し、AI エージェントのトークン消費コストを 52% 削減する成果を示した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月13日 22:06
AI深層分析
キーポイント
中国製オープンソースモデルの活用と再訓練
Writer は Palmyra X6 をゼロから学習させるのではなく、北京の Z.ai が開発した GLM-5.2 の重みを引き継ぎ、自社の米国インフラ上で独自にポストトレーニングを行った。
AI エージェントによるトークン消費爆発への対応
チャットボットと異なり AI エージェントは反復的な処理によりトークン消費が劇的に増加するため、Writer はコスト削減とガバナンス強化を最優先課題として解決策を提示した。
具体的な性能向上数値の発表
Palmyra X6 と新アーキテクチャの組み合わせにより、平均コストが 52% 低下し、速度は 48%、品質は 10% 改善されたことを同社が発表した。
企業における AI 導入の障壁変化
Writer の経営陣は、モデル能力そのものよりもコスト管理が現在の企業拡大における最大の障壁であると指摘し、予算抑制と採用拡大の両立を求めている。
合成データによるトレーニングとコスト削減
Palmyra X6 は教師モデルで生成した完全な合成データを基に訓練され、過去のモデルと同様に低コストで開発された。
重要な引用
"The biggest barrier today to enterprise expansion using AI is actually not model capabilities in most cases; it's actually the cost around them."
"This model is in no way, shape, or form connected to any of its original developers. It is fully run on our U.S. infrastructure."
"if an agentic task draws 20 times more tokens while unit prices fall 75%, total charges still rise fivefold."
"We do things like public benchmarks to let us know that we're climbing the right hill and that we don't have any sort of huge gaps, but we don't slavishly follow them either, because that's not really serving our customers."
編集コメントを表示
編集コメント
中国製モデルを基盤としながら米国インフラで再訓練する手法は、セキュリティとコストのバランスを取る新たなアプローチとして注目される。ただし、この戦略が長期的な技術的自立や地政学的リスク管理においてどの程度有効かは、今後の市場動向を見極める必要がある。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Fortune 500 企業であるアーンチュリー、Uber、バンガードなどが活用するエンタープライズ AI エージェントプラットフォーム「Writer」は本日、新フラッグシップモデル「Palmyra X6」を発表しました。これに合わせてエージェントのオーケストレーション基盤を再構築した「harness」と、突発的なトークン消費に対する IT リーダーのコントロールを可能にする新たなガバナンスツールも公開されています。
注目すべき数値があります。Writer によると、Palmyra X6 と組み合わせることで、エージェント製品の平均コストが 52% 削減され、処理速度は 48% 向上、品質も 10% 改善したとのことです。しかし、より重要なのは企業がこれを実現するためにどのような選択をしたか、そしてそれがエンタープライズ AI マーケットの行方をどう示唆しているかという点です。
Palmyra X6 はゼロから学習されたモデルではありません。北京に拠点を置く Z.ai(旧・Zhipu AI)が開発したオープンウェイトの混合専門家モデル「GLM-5.2」をポストトレーニングしたバージョンです。この事実は Writer が技術レポートで率直に開示しており、サンフランシスコにある同社は業界内で最も議論を呼ぶテーマの中心に立っています。それは、米国のエンタープライズが中国発のオープンソース基盤の上に構築すべきかどうかという問いです。
「このモデルは、元の開発者とは一切関係ありません。完全に米国のインフラ上で稼働しています」と、Writer のプロダクトマネジメント担当ディレクターである Matan-Paul Shetrit 氏は、発表前の独占インタビューで VentureBeat に語りました。
Writer の AI 研究を率いる Dan Bikel は、より率直にこう述べています。「これは明らかに Palmyra モデルです。私たちは単に浮動小数点の数値を取得してスタート地点とし、そこから学習を進めているだけです。」
なぜ AI エージェントはチャットボットとは異なる形で企業予算を圧迫しているのか
Writer の発表は、エージェント型 AI の経済性が企業の購買決定の中心に据えられたタイミングでなされました。通常、ユーザーのリクエストに対して 1 つの回答を生成するだけのチャットボットとは異なり、AI エージェントは単一のリクエストを、計画、情報取得、ツール呼び出し、検証、再試行といった反復的なループに変換します。このループのたびに計測されるトークンが消費されます。ユーザーに見えるのは 1 つの回答ですが、請求書にはその背後にある一連の処理全体が反映されるのです。
問題の規模は明らかになりつつあります。ゴールドマン・サックスの予測によると、2026 年から 2030 年の間にトークン消費量は 24 倍に増加し、月間 120 京(quadrillion)トークンに達する見込みです。これは質問をする人の数が増えたからではなく、常時稼働する企業用エージェントがその要因となっています。同分析では、トークン単価の下落が必ずしも請求額の減少を意味しないとも警告しています。例えば、エージェント型タスクで消費されるトークン量が 20 倍になり、単位価格が 75% 低下しても、総額は依然として 5 倍に増加する可能性があります。
「企業はトークン消費量の爆発的増加を歓迎します。それは採用が進んでいることを意味しますが、同時にコストは横ばいに抑えたいと考えています」と、Writer の CTO 兼共同創業者である Waseem AlShikh は述べています。
Writer の CEO、Shetrit は、企業における AI 導入の最大の障壁はモデルそのものの性能ではなく、コストにあると指摘しました。「現在の企業が AI を拡大する際の最大の障壁は、多くの場合、モデルの能力そのものではなく、それを取り巻くコストです」と彼は語りました。現実問題として、AI の代替手段が別の AI であるケースはほとんどなく、むしろ人手労働であることがほとんどです。
顧客のトークン使用量を削減することが、Writer 自身のトークン単価収益を損なうのではないかという質問に対し、Shetrit はその前提自体を否定しました。「コストを下げることは、私の利益に悪影響を与えるどころか、むしろ拡大させることになります。なぜなら、組織内での機会領域(TAM)が広がるからです」と彼は説明し、タスクあたりのコストを下げることで、これまで自動化されなかった業務フローが可能になると主張しました。この考え方は、クラウド時代の典型的なパターンと重なります。あの時代も単価は 10 年以上にわたり低下しましたが、利用量の拡大により請求総額は増え続けました。Writer は、エージェント分野でも同様の現象が起きると明確に見込んでおり、そこから利益を生み出すことに賭けているのです。
Palmyra X6 の内部:7440 億パラメータのモデルを 626 件のトレーニング例で微調整する方法
Writer の技術報告によると、Palmyra X6 は約 400 億のパラメータが各トークンで活性化する 7,440 億パラメータの混合専門家(MoE)モデルであり、GLM-5.2 のアーキテクチャをそのまま継承しています。同社が貢献したのは、意図的に保守的なポストトレーニング手法です。その名も「アンカー付き教師あり微調整(ASFT)」と呼ばれるこの技術は、わずか 626 件の厳選された合成エージェント軌跡という驚くほど小さなデータセットに対して、低学習率で単一のエポックのみを適用するものです。
この微小なデータセットこそが本質であり、制限事項ではありません。ASFT はトークン重み付けスキームと、KL 発散に基づく「アンカー」を組み合わせています。これは微調整後のモデルがベースモデルの凍結コピーから遠く逸脱しないようペナルティを与える仕組みで、新しいツール使用行動を教える一方で、ベースモデルが既に持つ汎用能力を損なうことなく学習を進めます。また Writer は、モデルのコア重み行列に対して標準的な Adam オプティマイザーに代わり、重み行列を幾何学的対象として扱う新しい手法である Muon を採用しました。
「『少ないほど良い』という哲学に従った論文が次々と出てきている」と Bikel 氏は述べ、小規模で極めて高品質なデータセットが非常に大きな効果をもたらすことを示した研究を引用しました。彼は続けて、「それが私たちがこのモデルを構築する際に遵循した哲学の一つであり、その結果も現れています。これにより、最適化作業のコストを下げつつ顧客に最適化できる体制を整えられ、最終的に私たちにとっても顧客にとっても低コストなモデルを実現できました」と語っています。
学習データはすべて合成データで構成されています。各プラン、ツール呼び出し、最終回答はいずれも教師モデルによって機械生成され、構造的な品質ゲート、モデルベースの検証器、そして 2 名の LLM からなる審査パネルによるフィルタリングを経てから学習に投入されます。これは Writer の長年の慣行を引き継ぐものです。同社の Palmyra X004 は 2024 年に約 70 万ドルでほぼ完全に合成データのみを用いて訓練され、TechCrunch がその当時報じていました。また SiliconANGLE によると、Palmyra X5 の学習には約 100 万 GPU 時間が必要でした。
Writer の内部評価では、グラウンディングや検索、ツール使用、コンテンツ生成、サブエージェントへの委任、ブランドボイスなど 9 つの能力において、X6 は平均 0.87(満点 1.00)を記録。Anthropic の Claude Opus 4.8(0.86)、Claude Sonnet 4.6(0.85)、OpenAI の GPT-5.5(0.80)、Google の Gemini 3.1(0.77)を上回りました。真に決定的な違いは価格にあります。Writer は X6 を、入力トークン 100 万あたり 2 ドル、出力トークン 100 万あたり 8 ドルで提供しています。一方、Opus 4.8 の料金はそれぞれ 15 ドルと 75 ドルです。同社によると、X6 はタスク完了に平均 26 秒を要し、単一の目標に向かって最大 8 時間、無人で作業を継続できます。
同社は内部ベンチマークに対して懐疑的な見方が寄せられることを率直に認めています。自社で評価した結果の公開を問われた際、Bikel氏は技術レポートには「公開ベンチマークの実施手順と、社内での評価結果の両方を記載している」と説明しました。彼は公開ベンチマークについて、「正しい方向に進んでいるか、大きな欠陥がないかを確認するための sanity check(健全性チェック)であり、盲目的に従うものではない」と述べています。「顧客にとって真に役立つのは、ベンチマークの数値を追求することではなく、実用的な成果を出すことだからです」。
中国に関する問い:GLM-5.2 をベースに構築することが企業セキュリティと信頼性にどう影響するか
Writer がベースモデルとして GLM-5.2 を選んだことは、2 年前であればアメリカのエンタープライズベンダーが考えることもできなかったことです。しかし現在、これは市場の実態を反映しています。今年 6 月に MIT ライセンスの下で公開された GLM-5.2 は、おそらく世界で最も能力の高いオープン利用可能なモデルです。独立分析機関 Artificial Analysis の知能指数では 51 点を獲得し、DeepSeek V4 Pro や Kimi K2.6 を上回り、Google の Gemini モデルの一部よりもエージェントタスクにおいて高い評価を得ています。また、欧州のテックメディア Trending Topics が報じたように、米国製フラッグシップ API の価格を大幅に下回るコストで提供されています。Writer のプレスリリースでは、「現在利用可能なオープンウェイトモデルの中で最も強力なモデル」と称しています。
オープンウェイトの急拡大には、本質的な課題も伴います。AI セーフティ非営利団体 SaferAI が 8 月に発表したレポートによると、GLM-5.2 は Z.ai の公開 API を通じて与えられた攻撃的なサイバーや生物学関連のタスクに対して拒否を示さず、Z.ai も安全性フレームワークや事前展開リスク評価を公表していませんでした。誰でも重みをダウンロードして改変できるようになれば、このギャップはさらに拡大します。
Writer 社の回答では、起源よりも「出所」の明確化とトレーニング後の調整が重要だと強調しています。ビケル氏は同社が「米国の Hugging Face から重みを取得し、米国インフラ上で完全にトレーニングを行った」と指摘。技術レポートによれば、すべてのデータセットは米国で合成・保存され、トレーニングに使用されたハードウェアもすべて米国国内に配置されていました。
同社はまた、政治的バイアス、検閲、事実性、拒否行動を網羅した、例外的に厳格な事前登録モデルリスク評価を実施しました。盲検審査員が 19,674 件の回答を採点し、X6 を GLM-5.2 ベースモデルおよび 4 つの最前線制御モデルと比較検証しています。
ワシントン・ポストの ModelSlant 政治バイアス評価では、Writer 社によると X6 はホットな話題について両論を提示する頻度が 80% に達し、これはテストされたすべてのモデルの中で最高水準でした。また、DeepSeek V4 が明確に拒否した政治的に敏感なプロンプトにも回答しています。FORTRESS の敵対的セーフティベンチマークでは、展開時のシステムメッセージを組み込んだ X6 は、生身の GLM-5.2 ベースモデルよりも敵対的安全性で 8.6 ポイント高く評価されつつも、有用性の低下はほとんどありませんでした。
「バイアスや検閲に関する広範なベンチマークを実施しました」とシェトリ氏は語り、「ダン氏とチームが取り組んだ成果により、このモデルはオープンソースの代替品だけでなく、ベンチマーク実施時点で市場にあったクローズドソースの代替品よりもはるかに優れていることが実証されました。」
同レポートには重要な留保事項も含まれています。英語での振る舞いにおいて統計的に有意な政治的な非対称性は確認されませんでしたが、「言語によって振る舞いが異なることが示された」と指摘されています。これは、626 回のファインチューニングの試行がベースモデルの訓練痕跡をすべて消し去れるわけではないという率直な認めに他なりません。
ハネス効果:オーケストレーションこそがモデルそのものより重要かもしれない
Writer の発表の中で最も戦略的に興味深い主張は、実は Palmyra X6 自体には一切関係ありません。同社が再構築した Writer Agent ハネス(タスクの計画、作業の一括処理、サブエージェントへの委任、コンテキスト管理を行うオーケストレーション層)は、Anthropic や OpenAI からのサードパーティ製モデルを含むあらゆるテスト対象モデルで、コストを 41% 削減し、タスク完了時間を 44% 短縮しながら品質も維持できると述べています。この「ハネス効果」に関する知見は、同社が発表に併せて公開した研究論文の中で詳しく解説されています。
これにより当然ながら疑問が生じます。VentureBeat が同社へ投げかけた質問のように、「ハネス単体ですべてのモデルで最大の節約を実現できるなら、なぜわざわざモデルを構築する必要があるのか?」という点です。
シェトリ氏の回答は「制御」に関するものでした。「実験室がモデルを非推奨にするかどうか、またそのモデルでどのようなデータを使用するかについては、私にはコントロールできません」と彼は語りました。「しかし、私がモデルを構築する際には、大きな道徳的責任を負うことになります。そして、企業顧客から寄せられる困難な質問にもお答えできます。」
ビケル氏は、このモデルとハネス(制御枠組み)は一体となって開発されたと付け加えました。「このモデルはハネスと共に、本質的に共進化してきました。私たちは『Writer Agent』という旗艦製品を持っています。この製品をこのモデルと完璧に連携させたいと考えており、実際にその通りになっています。モデル開発の過程でもこれを考慮しており、自社でモデルを構築しない限り実現できないことです。」
值得注意的是、Writer は同時にリスクヘッジも図っています。今回のリリースにより、同社は Writer Agent におけるマルチモデルサポートを拡大し、管理者が Anthropic や OpenAI のモデル、さらに Microsoft Azure、AWS Bedrock、Nvidia NIM といったクラウドプロバイダーのモデル、さらには Writer が自社で構築していない画像生成モデルさえも利用可能にしました。CIO(最高情報責任者)に向けたメッセージは極めてシンプルです。「当社のモデルを使用すれば、ワークフローにとって最も安価かつ最適ですが、プラットフォームを利用する限り、どちらを選んでもコスト削減につながります。」
新しいガバナンスツールは、予期せぬ AI 請求書の発生を未然に防ぐことを目指しています
今回のリリースの3つ目の柱は、より静かながらも深刻な企業の課題に焦点を当てています。それは経営陣がAIエージェントの支出内容を把握できていないという問題です。
新しいガバナンスツールにより、管理者は企業全体のエージェント利用状況を一元管理できるようになります。また、社内で共有可能な「プレイブック」や「スキル」自動化ワークフローごとの分析機能も提供され、アラート通知や支出上限を設定できる消費量制御機能も備えています。
支出制限の導入が顧客に予期せぬ請求が発生したことを意味するのかと問われた際、シェトリ氏はこれを被害拡大防止ではなく、採用を促進するものとして捉え直しました。「CIOやCISOとして、セキュリティ面とコスト面の両面で安心感を持っていただき、組織内でAI利用を拡大できるようにするためのツールをどう構築するか。それが私たちの目指すところです」と彼は語りました。
彼の説明によれば、可視性が経営層の承認を得る鍵となります。明確なコストデータを持つ企業は「実際には、これまで手をつけたことのないユースケースにもAI導入を拡大しようと考えている」とのことです。
この機能群は、企業がAIに予算を配分する方法における広範な転換を反映しています。フォーブスがトークン価格競争について分析したように、洗練された購入者は、広告されているレートカードに期待される呼び出し回数を単純に乗算するのではなく、成功したタスクあたりのコストをモデル化するようになっています。具体的には、リトライ回数やツール呼び出し、エスカレーションなどをカウントして計算します。
Writerはこの専門的な考え方を製品化し、これまで財務チームがスプレッドシートで行っていた作業を、ネイティブなプラットフォーム機能へと昇華させています。
同社は、過去1年以上かけて構築してきたガバナンスの枠組みも今回で完結させました。Writerは昨年11月に管理者による制御機能を備えた統一型エージェント体験をリリースし、今年3月にはエージェントのスキル機能やワークフロー分析機能を追加しました。今回の木曜日の発表では、すべてのワークフローに価格設定と支出上限を設定することで、このガバナンスサイクルが完結します。
Writerは2020年にMay Habib氏とWaseem AlShikh氏によって設立され、2024年末には19億ドルの評価額で2億ドルの資金調達を果たしました。同社のビジネスモデルは、消費者向けの大規模展開ではなく、規制が厳しくリスクの高いエンタープライズ環境での導入に特化しています。創業者のShetrit氏は、この狭い焦点に対する謝罪は一切していません。「エンタープライズのユースケースに集中して取り組む特権とは、自社のモデルでフランス語のソネット詩を書ける必要がないということです」と彼は述べました。「何でもやろうとしないからこそ、顧客が抱える課題やニーズに集中できるのです」。
同社が何者であるかについても、彼は率直な見解を述べています。「私たちは研究機関から転身して消費者向け製品を手掛けた企業が、たまたまエンタープライズ市場にも参入しているわけではありません。私たちは第一に、エンタープライズ顧客のために活動する企業です。すべての意思決定は、この視点に基づいて評価されます。つまり、ゼロから構築することが最適だと判断すれば、そうします。一方で、市場には自社の顧客により良いサービスを提供できる代替手段があると考えれば、それらを活用することも厭いません」。
この実用主義こそが、今回のリリースが示す最も重要なシグナルかもしれません。十分な資金を有し、5 年間にわたるモデル構築の経験を持つ米国の AI 企業が、価値の最前線はもはや事前学習にはないと結論付けました。重要なのは最後の 1 マイル、つまりオープンウェイトのポストトレーニング、その周囲を囲むエンジニアリング、そして CFO にダッシュボードを手渡すことです。
Writer の見方が正しければ、最先端研究所の競争優位性は、品質が実際に 7 倍の価格プレミアムを正当化できるワークロードに限定されます。それ以外の領域では、誰かが事前学習のために支払ったモデルこそが勝者となります。
3 年間も「どのモデルが最も賢いか」で議論が続く業界において、Writer は異なる賭けに出ています。エンタープライズ AI の競争は、浮動小数点演算性能の数値が最も優れた企業によって勝利されるのではなく、その数値をどう活用するかを知っている企業が勝つのです。
原文を表示
Writer, the enterprise AI agent platform used by Fortune 500 companies including Accenture, Uber, and Vanguard, released its new flagship model Palmyra X6 today, alongside a rebuilt agent orchestration "harness" and new governance tools designed to give IT leaders control over runaway token spending.
The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6. But the more consequential story may be how the company got there — and what its choices reveal about where the enterprise AI market is heading.
Palmyra X6 is not trained from scratch. It is a post-trained version of GLM-5.2, the open-weight mixture-of-experts model from Beijing-based Z.ai, formerly Zhipu AI — a fact Writer discloses openly in its technical report, and one that places the San Francisco company at the center of one of the industry's most charged debates: whether American enterprises should build on Chinese open-source foundations.
"This model is in no way, shape, or form connected to any of its original developers. It is fully run on our U.S. infrastructure," Matan-Paul Shetrit, Writer's director of product management, told VentureBeat in an exclusive interview ahead of the announcement.
Dan Bikel, who leads Writer's AI research, put it more bluntly: "It's very much a Palmyra model, and we just happen to grab the floating point numbers as the starting point, and train from there."
Why AI agents are blowing up enterprise budgets in ways chatbots never did
Writer's announcement lands at a moment when the economics of agentic AI have moved to the center of enterprise buying decisions. Unlike a chatbot, which typically generates one answer per user request, an AI agent turns a single request into repeated rounds of planning, retrieval, tool calls, validation, and retries — with every loop consuming metered tokens. The user sees one answer; the invoice reflects the entire loop.
The scale of the problem is becoming clear. Goldman Sachs forecasts that token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month, driven not by more people asking questions but by always-on enterprise agents. The same analysis warned that falling per-token prices do not guarantee falling bills: if an agentic task draws 20 times more tokens while unit prices fall 75%, total charges still rise fivefold.
"The enterprise wants token consumption to explode — it means adoption is happening — but they need costs to flatten," said Waseem AlShikh, Writer's CTO and co-founder, in a statement.
Shetrit framed the cost problem as the primary obstacle to enterprise AI adoption — more so than model capability itself. "The biggest barrier today to enterprise expansion using AI is actually not model capabilities in most cases; it's actually the cost around them," he said. "The reality today is, in most cases, the alternative for AI is not another AI, it is human labor."
Asked whether cutting customers' token consumption would cannibalize Writer's own per-token revenue, Shetrit rejected the premise. "Reducing the cost is not hurting my bottom line. It's actually expanding it, because it's expanding the TAM of opportunity within an organization," he said, arguing that lower per-task costs unlock workflows enterprises would otherwise never automate. That argument echoes a pattern familiar from the cloud era, where unit prices fell for a decade while total bills rose as consumption expanded — a dynamic Writer is explicitly betting will repeat with agents, and betting it can profit from.
Inside Palmyra X6: how 626 training examples fine-tuned a 744-billion-parameter model
Palmyra X6 is a 744-billion-parameter mixture-of-experts model with roughly 40 billion active parameters per token, inheriting GLM-5.2's architecture unchanged, according to Writer's technical report. The company's contribution is a deliberately conservative post-training recipe: a technique called anchored supervised fine-tuning (ASFT), applied to a remarkably small corpus of just 626 curated synthetic agentic trajectories, trained for a single epoch at a low learning rate.
The tiny dataset is the point, not a limitation. ASFT pairs a token-weighting scheme with a KL-divergence "anchor" that penalizes the fine-tuned model for drifting too far from a frozen copy of the base model — teaching new tool-use behaviors without eroding the general capabilities the base already has. Writer also swapped the standard Adam optimizer for Muon, a newer method that treats weight matrices as geometric objects, on the model's core weight matrices.
"There's a whole string of papers following a quote-unquote 'less is more'" philosophy, Bikel said, referencing research showing that "small, extremely high quality data sets go a really long way." He added: "That's the philosophy — one of the philosophies — that we followed when building this model, and it showed. It allowed us to optimize for our customers at lower cost to do the work of optimization, and that ultimately yielded a lower cost model for us and for them."
The training data itself is fully synthetic — every plan, tool call, and final answer machine-generated by teacher models, then filtered through structural quality gates, a model-based verifier, and a two-model LLM judging panel before entering training. That continues a long-standing Writer practice: the company's Palmyra X 004 was trained almost entirely on synthetic data for roughly $700,000 back in 2024, as TechCrunch reporte at the time, and Palmyra X5 required about $1 million in GPU hours, according to SiliconANGLE.
On Writer's internal evaluations — nine capabilities spanning grounding and retrieval, tool use, content generation, sub-agent delegation, and brand voice — X6 scored an average of 0.87 out of 1.00, edging out Anthropic's Claude Opus 4.8 (0.86), Claude Sonnet 4.6 (0.85), OpenAI's GPT-5.5 (0.80), and Google's Gemini 3.1 (0.77). The price gap is the real differentiator: Writer prices X6 at $2 per million input tokens and $8 per million output tokens, versus 15/75 for Opus 4.8. The company says X6 completes tasks in 26 seconds on average and can work unattended toward a single goal for up to eight hours.
Writer is candid that internal benchmarks invite skepticism. Asked directly whether the company would publish its methodology after grading its own homework, Bikel said the technical report covers "both the protocol we used to do our public benchmarking as well as our internal evaluations." He described public benchmarks as sanity checks rather than targets: "We do things like public benchmarks to let us know that we're climbing the right hill and that we don't have any sort of huge gaps, but we don't slavishly follow them either, because that's not really serving our customers."
The China question: what building on GLM-5.2 means for enterprise security and trust
Writer's choice of base model would have been unthinkable for an American enterprise vendor two years ago. Today it reflects a market reality: GLM-5.2, released in June under the permissive MIT license, is arguably the most capable openly available model in the world. Independent analysis house Artificial Analysis scored it at 51 on its Intelligence Index — ahead of DeepSeek V4 Pro, Kimi K2.6, and even some of Google's Gemini models on agentic tasks — while undercutting U.S. flagship API pricing many times over, as European tech outlet Trending Topics reported. Writer's press release calls it "the strongest available open-weight model."
The open-weight surge carries genuine baggage. An August report from AI safety nonprofit SaferAI found that GLM-5.2 refused none of the offensive cyber or biology tasks it was given via Z.ai's public API, and that Z.ai published no safety framework or pre-deployment risk assessment — a gap that widens once anyone can download and modify the weights.
Writer's answer is that provenance and post-training matter more than origin. Bikel emphasized that the company "grabbed the weights off of the U.S. Hugging Face" and trained entirely on American infrastructure; the technical report states all datasets were synthesized and stored in the U.S., and all training hardware was located in the U.S.
The company also ran what it describes as an unusually rigorous, pre-registered model-risk evaluation covering political bias, censorship, factuality, and refusal behavior — 19,674 evaluated responses scored by blinded judges — comparing X6 against its GLM-5.2 base and four frontier control models.
On the Washington Post's ModelSlant political-bias evaluation, Writer says X6 presented both sides of hot-button questions 80% of the time, the highest rate of any model tested, and answered politically sensitive prompts that DeepSeek V4 refused outright. On the FORTRESS adversarial safety benchmark, X6 with its deployment system message scored 8.6 points higher on adversarial safety than the raw GLM-5.2 base, at negligible cost to benign helpfulness.
"We've run extensive benchmarking around bias, around censorship," Shetrit said, "and the work Dan and the team has done has actually proven that this model is actually significantly better than not just open source alternatives, but any closed source alternative in the market at the time of the benchmarking."
The report does hedge in one notable place: while English-language behavior showed no statistically robust political asymmetry, "the behavior was shown to vary by language" — a candid admission that 626 fine-tuning trajectories do not scrub every trace of a base model's training.
The harness effect: why orchestration may matter more than the model itself
Perhaps the most strategically interesting claim in Writer's announcement has nothing to do with Palmyra X6 at all. The company says its rebuilt Writer Agent harness — the orchestration layer that plans tasks, batches work, delegates to sub-agents, and manages context — cuts costs by 41% and completes tasks 44% faster across every model it tested, including third-party models from Anthropic and OpenAI, while maintaining quality. Writer published the finding in an accompanying research paper on what it calls "The Harness Effect."
That raises an obvious question, which VentureBeat put to the company: if the harness alone delivers most of the savings on any model, why build a model at all?
Shetrit's answer was about control. "I cannot control if a lab deprecates their model. I cannot control what data they use in their model," he said. "Where when I build the model, I have significant moral control, and I can answer the tough questions that enterprise customers ask me."
Bikel added that the model and harness were developed together: "This model was built and essentially co-evolved with the harness... We know that we have a flagship product, Writer Agent. We want that to work really, really well with this model, and sure enough, it does. And we take that into account during model development, and that's something that is not possible if you don't build your own model."
Notably, Writer is simultaneously hedging. With this release, the company extends multi-model support to Writer Agent, letting admins enable models from Anthropic, OpenAI, and cloud providers including Microsoft Azure, AWS Bedrock, and Nvidia NIM — even image-generation models, a category Writer does not build. The message to CIOs is disarmingly simple: use our model because it is cheapest and best for your workflows, but the platform saves you money either way.
New governance tools aim to end surprise AI bills before they start
The third leg of the release targets a quieter enterprise pain point: nobody in the C-suite knows what the agents are spending. New governance tools give administrators a centralized view of agent usage across the business, per-workflow analytics for the company's shareable "Playbooks" and "Skills" automations, and consumption controls with alerts and spending limits.
Asked whether the introduction of spending controls implied that customers had been receiving surprise bills, Shetrit reframed it as an adoption enabler rather than damage control. "How do we build the tools to allow you as the CIO, CISO in a company, to feel comfortable both on the security and spend, so you can expand AI usage in your organization," he said. In his telling, visibility is what lets leaders say yes: businesses with clear cost data "are actually looking to expand AI adoption to use cases that they would never have touched before."
The feature set tracks a broader shift in how enterprises budget for AI. As Forbes analysis of the token price wars argued, sophisticated buyers are learning to model cost per successful task — counting retries, tool calls, and escalations — rather than multiplying expected calls by the advertised rate card. Writer is effectively productizing that discipline, turning what has been a finance-team spreadsheet exercise into a native platform capability.
It also completes a governance arc the company has been building for over a year. Writer shipped its unified agent experience with admin controls last November, then added agent Skills and workflow analytics in March, according to earlier company announcements. Thursday's release closes the loop by attaching a price tag — and a spending limit — to every workflow.
Writer, founded in 2020 by May Habib and Waseem AlShikh, raised $200 million at a $1.9 billion valuation in late 2024, and has built its business on regulated, high-stakes deployments rather than consumer scale. Shetrit made no apology for the narrowness of that focus. "The privilege of working and focusing on enterprise use cases is that I don't need my model to be able to write a French sonnet," he said. "When you don't try to do everything, you can focus on your customer problem and needs."
He was equally direct about identity: "We are not a research lab converted to a consumer product now dabbling in enterprise. We are first and foremost an enterprise company that serves enterprise customers, and we evaluate our decisions within that lens. Which means, if we think building things from scratch is the right decision, that's what we will do. But if we think there are other alternatives out there in the market that serve our customers better, that's what we will do."
That pragmatism may be the release's most important signal. A well-capitalized American AI company with five years of model-building experience has concluded that the frontier of value no longer lies in pretraining, but in the last mile: post-training open weights, engineering the harness around them, and handing the CFO a dashboard. If Writer is right, the frontier labs' moat narrows to the workloads where quality genuinely justifies a sevenfold price premium — and for everything else, the winning model is the one somebody else paid to pretrain.
In an industry that has spent three years arguing about whose model is smartest, Writer is making a different wager: the enterprise AI race won't be won by the company with the best floating point numbers, but by the one that knows what to do with them.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み