Pokee AI、顧客境界内動作の 10M トークンエージェントモデル「Pokee-Isaac 28B」を公開
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
Pokee AI は、1000万トークンのコンテキストウィンドウを持つ「Pokee-Isaac 28B」を発表し、データが外部境界を越えない環境での長期エージェント運用を可能にする。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月9日 02:25
AI深層分析
キーポイント
10M トークン・コンテキストとオンプレミス対応
Pokee AI は 28B パラメータのテキスト専用モデル「Pokee-Isaac」をリリースし、1000万トークンのコンテキストウィンドウを実現した。このモデルは VPC、オンプレミス、またはデバイス上で実行可能であり、データが外部境界を越えない規制業界向けに設計されている。
ベンチマークにおける競合との同等性能
同社によると、Isaac は 10M トークンで RULER ベンチマークの 93.3% を達成し、コスト最適化されたクラウドベースラインと同等のパフォーマンスを示す。また、BFCL v4 や τ³-bench などのエージェントベンチでも Gemini 系列モデルと競合する結果を記録している。
単一 GPU での推論とライセンス形態
Isaac は OpenAI 互換の API を通じて提供され、RTX 4090 相当の単一 GPU で動作可能である。ただしオープンウェイトではなくライセンス販売であり、開発者向けには vLLM や SGLang の Day-0 サポートが謳われている。
10M トークンコンテキストでの卓越した性能と効率性
Pokee-Isaac 28B は 10M トークンのコンテキストで RULER ベンチマークに 93.3% のスコアを記録し、競合モデルは 2M トークンを超えると 0.0 に落ちる中、単一の B200 GPU で最大 137,200 トークン/秒のプレフィルスループットを実現する。
ベンチマークでの高い評価とセキュリティ強化
BFCL v4 と τ³-bench で他モデルを上回り、DTAP レッドチームテストでは最も低い攻撃成功率(35.6%)を記録しながらも 82.5% の良性タスク成功率を維持する。
重要な引用
Pokee AI released Pokee-Isaac 28B, a 28B text-only foundation model with a 10M-token context window, designed to run inside that boundary.
The Pokee research team claims 93.3% on RULER at 10M tokens, parity with the strongest cost-optimized cloud baselines on agentic benchmarks
when enough usable context is available in-boundary, memory hierarchies and compression become optional rather than required.
On DTAP red-teaming, Isaac records the lowest direct (36.0), indirect (35.2), and combined (35.6) attack success rates, with 82.5 benign success.
編集コメントを表示
編集コメント
10M トークンというコンテキスト規模は、実務における長文ドキュメント解析や複雑なエージェントタスクの解決において大きな転換点となる。ただし、ベンチマーク結果が高性能 GPU で取得されたものである点は留意が必要であり、実際の導入環境での性能検証が求められる。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
長期のホライズンを持つエージェントは、タスクを解決する速度よりもコンテキストを蓄積する速度が速くなります。ツールの出力、観測結果、中間推論のステップすべてがウィンドウ内に残り、「そのコンテキストを保持し」「それを通じて一貫性を保つ」という 2 つの重要な機能は、これまでほぼクラウドエンドポイントからしか利用できませんでした。しかし、データが境界を越えて流出することが許されない規制産業、公共機関、オンデバイスアプリケーションではこの制限が大きな課題となります。
Pokee AI は、この境界内で動作するように設計された 10M トークンコンテキストウィンドウを持つ 28B パラメータのテキスト専用ファウンデーションモデル「Pokee-Isaac 28B」をリリースしました。研究チームによると、10M トークンでの RULER ベンチマークで 93.3% のスコアを達成し、エージェント向けベンチマークではコスト最適化された最強のクラウドベースラインと同等のパフォーマンスを示しています。さらに、単一の GPU で動作可能なサービングプロファイルを実現しています。
導入は可能でしょうか?
はい、可能です。ただし、オープンウェイトではなくライセンス契約による提供となります。Pokee AI は Isaac を OpenAI 互換の開発者 API 経由で提供しており、VPC 内、オンプレミス、またはオンデバイスへのデプロイをライセンス契約で許可しています。
ローンチ発表では、vLLM と SGLang の Day-0 サポートと、RTX 4090 または同等の GPU から開始できる単一 GPU サービングが謳われています。ただし研究チームが公開している測定値は B200 クラスの GPU 1 基のみでの結果であるため、消費者向け GPU での動作に関する主張は、報告された結果というよりはベンダーからのガイダンスとして捉えるべきです。
企業レベルでの適用: このモデルは、推論スタックを自社で保有している組織に適しています。具体的にはプラットフォームチームを持つ中堅・大企業や、デバイス OEM 企業が対象です。オンプレミス環境のハードウェアを持たない個人利用者は、ホスト型 API の利用を検討すべきです。「境界内(インバウンダリー)」での運用メリットが活きるのは、明確なデータ境界が存在する場合に限られます。
適用業界: ヘルスケアと保険者、金融サービス、防衛・公共部門、法務および電子開示、製薬や半導体の研究開発などです。これらの業界に共通するのは、「データを外部 API の境界を越えて移動させてはならない」というルールが存在することであり、単なるプライバシーへの配慮ではありません。
適用事例: 全リポジトリを対象としたコードレビュー、数年単位のスパンを持つ契約書や請求書の分析、完全なログアーカイブに基づくインシデントのフォレンジック調査、要約やコンテキストの剪定を一切必要としない長時間稼働するツールエージェントなどです。研究論文ではこの点について明確に指摘されています。利用可能なコンテキストが境界内で十分に確保できる場合、メモリ階層や圧縮技術は必須ではなく、オプションとなるのです。
長文コンテキストにおける結果
RULER ベンチマークにおいて、Isaac はテストされたすべての長さで 93.3% を維持し、10M トークンの最終地点でも同スコアを達成しました。一方、GPT-5.6 Luna や Gemini 3.5 Flash Lite は 512K トークンまで追従しますが、1M トークンに達するとコンテキストオーバーフローを起こします。
8 つの「針(ニードル)」を含む MRCR v2 ベンチでは、Isaac は 256K、512K、1M の各長さでそれぞれ 0.607、0.743、0.500 のスコアを記録しました。この範囲全体を通じて、Gemini を上回る差は 0.133 から 0.295 に拡大しています。
エージェント機能とセキュリティにおける結果
Isaac は BFCL v4 で 70.94 のスコアを記録し、Luna の 70.61 をわずかに上回っています。しかし報告書はこの差を「リード」と呼ぶのではなく「互角」であると評価しており、この見解が妥当です。
τ³-bench では 4 つのドメインで平均 0.662 を達成し、Gemini の 0.631 を上回っています。特に銀行関連タスクでは難易度が高く、全モデルともスコアは 0.186 に留まりました。
MCP-Atlas ではカバレッジ率 74.59% で 3 位ですが、1 タスクあたりのターン数は 9.10 と Gemini の 14.99 よりも少ないです。Terminal-Bench 2.1 では、86 問中テキスト対応タスクの 56 問(65.1%)を解決し、Luna が 60 問だったことには及びません。これはクラウドベースラインが唯一勝利したベンチマークであり、報告書もその事実を率直に述べています。
DTAP のレッドチームテストでは、Isaac は直接攻撃(36.0)、間接攻撃(35.2)、そして総合的な攻撃成功率(35.6)で最も低い値を記録しました。一方で、良性タスクの成功率は 82.5% に達しています。ただし注意すべき点として、ベースラインモデルは標準ランナーで実行されたのに対し、Isaac は Pokee ハーネス上で評価されています。
効率性、価格、そしてポータビリティ
B200 クラスの GPU 1 基での RULER ワークロードでは、コンテキスト長 1M で TTFT(Time to First Token)は 23.6 秒、10M では 72.9 秒でした。プリフィル処理のスループットはコンテキスト長とともに向上し、42,400 トークン/秒から 137,200 トークン/秒に達します。つまり、プロンプトが 10 倍長くても、TTFT はおよそ 3 倍の増加で済みます。
リスト価格は入力・出力ともに 100 万トークンあたり $0.15/$1.00 とされています(暫定価格)。さらに、Isaac は Intel Arc Pro B70 や Core Ultra Series 3 (Panther Lake)、Qualcomm Snapdragon X2 Elite 上でフルデバイス動作が可能です。
キーポイント
- Pokee-Isaac 28B は 10M トークンで RULER に 93.3% のスコアを記録し、同パネルのすべてのベースラインモデルは 2M を超えるとスコアが 0.0 になります。
- B200 1 基上で 10M コンテキスト時のプリフィル処理スループットは 137,200 トークン/秒に達し、デコード速度は約 335 トークン/秒で一定を維持します。
BFCL v4 で 70.94、τ³-bench では平均 0.662 を記録し、それぞれ首位を獲得。Terminal-Bench 2.1 では 2 位、MCP-Atlas では 3 位という結果を残しています。
DTAP における攻撃成功率は全体で最低の 35.6% に抑えつつ、通常のタスクの成功率は 82.5% を維持。高い安全性と実用性を両立させています。
モデルの重み(Weights)は公開されていません。VPC 内、オンプレミス環境、またはデバイス上でのみライセンス契約に基づき展開可能です。
詳細はブログや論文をご覧ください。Twitter でフォローいただくこともできますし、15 万人以上の ML 関係者が集まる SubReddit やニュースレターへの登録もぜひご検討ください。Telegram をご利用の方も、今ならこちらからも参加できるようになりました。
本記事は MarkTechPost にて公開された「Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary」の翻訳です。
原文を表示
Long-horizon agents accumulate context faster than they resolve tasks. Every tool output, observation, and intermediate reasoning step stays in the window, and the two capabilities that matter — holding that context and staying coherent across it — have so far been available almost exclusively from cloud endpoints. That excludes regulated industries, public-sector institutions, and on-device applications, where the data is not permitted to leave the boundary at all. Pokee AI released Pokee-Isaac 28B, a 28B text-only foundation model with a 10M-token context window, designed to run inside that boundary. The Pokee research team claims 93.3% on RULER at 10M tokens, parity with the strongest cost-optimized cloud baselines on agentic benchmarks, and a serving profile that fits a single GPU.
Is it deployable
Yes — but licensed, not open-weight. Pokee AI serves Isaac through an OpenAI-compatible developer API, and licenses it for deployment inside a VPC, on-premises, or on-device. The launch announcement advertises Day-0 support for vLLM and SGLang, and single-GPU serving starting from an RTX 4090 or equivalent. The research team publishes measurements only from a single B200-class GPU, so treat the consumer-GPU claim as vendor guidance rather than a reported result.
Company level: This fits organizations that already own their inference stack — mid-size and enterprise teams with a platform group, plus device OEMs. A solo practitioner without on-prem hardware should use the hosted API instead; the boundary argument only pays off if you have a boundary.
Industries: Healthcare and payors, financial services and insurance, defense and public sector, legal and e-discovery, and pharma or semiconductor R&D. The common trait is a rule that the data cannot cross an external API boundary, not a preference for privacy.
Applications: Whole-repository code review, multi-year contract and claims analysis, incident forensics over full log archives, and long-running tool agents that never need summarization or context pruning. The research paper makes this second point explicitly: when enough usable context is available in-boundary, memory hierarchies and compression become optional rather than required.
Long-context results
On RULER, Isaac stays above 93.3% at every tested length, ending at 93.3% at 10M. GPT-5.6 Luna and Gemini 3.5 Flash Lite track it to 512K, then hit context-overflow at 1M.
On MRCR v2 with 8 needles, Isaac scores 0.607, 0.743, and 0.500 at 256K, 512K, and 1M. Its margin over Gemini widens from 0.133 to 0.295 across that sweep.
Agentic and security results
Isaac leads BFCL v4 at 70.94 against Luna’s 70.61. The report calls that parity rather than a lead, which is the correct read. On τ³-bench it averages 0.662 across four domains, ahead of Gemini’s 0.631, with banking at 0.186 for everyone’s difficulty. On MCP-Atlas it places third at 74.59% coverage, but uses 9.10 turns per task against Gemini’s 14.99. On Terminal-Bench 2.1 it resolves 56 of 86 text-compatible tasks (65.1%), behind Luna’s 60. That is the one benchmark a cloud baseline wins, and the report states it plainly.
On DTAP red-teaming, Isaac records the lowest direct (36.0), indirect (35.2), and combined (35.6) attack success rates, with 82.5 benign success. One condition differs: baselines ran under the stock runner, Isaac under the Pokee harness.
Efficiency, pricing, and portability
Under the RULER workload on one B200-class GPU, TTFT is 23.6s at 1M and 72.9s at 10M. Prefill throughput rises with context, from 42,400 to 137,200 tokens/s, so a ten-fold longer prompt costs roughly three times the TTFT. List pricing is $0.15/$1.00 per million input/output tokens, marked provisional. Isaac also runs fully on-device on Intel Arc Pro B70 and Core Ultra Series 3 (Panther Lake), and on Qualcomm Snapdragon X2 Elite.
Key Takeaways
Pokee-Isaac 28B scores 93.3% on RULER at 10M tokens; every baseline in its panel returns 0.0 beyond 2M.
Prefill reaches 137,200 tokens/s at 10M context on one B200; decode holds flat near 335 tokens/s.
It leads BFCL v4 (70.94) and τ³-bench (0.662 avg), places second on Terminal-Bench 2.1, third on MCP-Atlas.
Lowest combined attack success rate on DTAP (35.6) while keeping 82.5 benign task success.
Weights are not published; deployment is licensed into VPC, on-premises, or on-device.
Check out the Blog and Paper. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agentic Model Built to Run Inside the Customer Boundary appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み