フランスのスタートアップ Kog、既存 GPU で推論速度向上を目指す
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TechCrunch AI
フランスのスタートアップ Kog は既存のデータセンター向け GPU で推論速度を劇的に向上させるソフトウェア最適化技術を開発し、企業顧客からの関心を集めている。
AI深層分析を開く2026年8月15日 01:22
AI深層分析
キーポイント
既存ハードウェアでの高速化戦略
Kog は独自チップに依存せず、AMD MI300X や Nvidia H200 などの既存のデータセンター向け GPU で推論速度を最大化するソフトウェア最適化技術に注力している。
市場からの具体的な反応
CEO Gaël Delalleau によると、同社の発表により 200 の実質的なビジネスリードを獲得しており、推論速度とコストがボトルネックとなっている現状への需要が高いことを示している。
主要なターゲットユースケース
初期フィードバックに基づき、開発者向けのコード生成ツールや、プロンプトでゲーム・アプリを生成するデザインパートナー企業を主な顧客層として想定している。
大規模モデルへの集中
顧客が小規模モデルのファインチューニングに準備できていないという知見から、Kog は現在、より大きな大規模言語モデル(LLM)の加速開発に全焦点を当てている。
GPUの非効率性に対する認識の転換
CEOはLLM推論にGPUが不適切という考えを誤解とし、新しいGPUのメモリ帯域幅を活用する可能性を強調している。
重要な引用
"We had 200 tangible business leads," CEO Gaël Delalleau told TechCrunch.
"And that's why since the launch, we've been fully focused on accelerating the development of larger models to meet the demand we've seen."
"GPUs have a bright future," he said.
"to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn't necessarily designed."
編集コメントを表示
編集コメント
Kog のアプローチは、AI インフラの民主化という観点から極めて意義深い。専用ハードウェアへの巨額投資を回避しつつ、既存資産で高性能を発揮できる点は、多くの企業が直面するコスト課題に対する即効性のある解決策となり得る。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
AI推論の高速化を巡る競争は激化しており、市場は Cerebras とその専用チップに 5 月の IPO 初日で 温かい歓迎 を送りました。しかし、フランスのスタートアップ「Kog」は、既存の汎用 GPU からさらに多くの性能を引き出す余地があると信じています。
同社は 5 月、Hacker News のトップページに登場し、企業がすでに保有する標準的なデータセンター向け GPU(デモでは AMD MI300X や Nvidia H200 を使用)でも「極めて高速な単一リクエストのデコードが可能である」ことを示す 技術プレビュー を発表しました。
一部の人は、この成果がラップトップ搭載の GPU には適用できないことに失望しましたが、他の人々はその可能性に注目しました。推論速度とコストが現在、決定的なボトルネックとなっている中、既存ハードウェアでソフトウェア最適化により新たな能力を引き出すという Kog の約束は、単なる傍観者を超えて 多くの注目を集めています。「200 件の具体的なビジネスリードを獲得しました」と CEO の Gaël Delalleau は TechCrunch に語っています。
初期フィードバックに基づき、創業者のソロ氏はソフトウェアエンジニアリングを最初のユースケースとして想定しています。Claude Code のベテランユーザーはよくご存知の通り、結果を得るまでに数時間待たされることもあります。Anthropic 社自身も「速度はお金」と認識しており、Claude の Fast Mode には 価格の倍数 を請求しています。
Kog はこうした遅延に辟易している顧客をターゲットにしようとしています。彼らは通常、専門的な業務で AI ワークフローに依存しています。しかし同社は、プロンプトだけでゲームやアプリを生成できるデザインパートナーも抱えており、Kog Inference Engine(KIE)による高速化が収益増につながるという点でも意義があると、Delalleau 氏は語っています。
同社は、この市場はまだ成熟していないことを理解しています。需要を観察する中で、潜在的な顧客は小規模モデルのファインチューニングには準備できていないことがわかりました。「そのため、ローンチ以来、私たちが目にした需要に応えるため、大規模モデルの開発加速に完全に注力してきました」と述べています。
これにより、「30 倍高速な LLM 推論」という約束を実現するには大きな飛躍が必要です。デモでは印象的な 1 リクエストあたり 3,000 トークン/秒(TPS)を達成しましたが、これは約 20 億パラメータのみに特化した専用小規模モデル Laneformer 2B を使用した結果です。同モデルは現在 オープンソース化されています。
懐疑論者たちをよそに、デラローは同様のアプローチが LLM(大規模言語モデル)にも十分に機能すると確信しています。LLM のサイズは推論チップにとって課題となることもありますが、「GPU には明るい未来がある」と彼は断言します。Kog の CEO にとっても、GPU がデコードに適していないという考えは誤解であり、最新の GPU はますます増大するメモリ帯域幅を有しており、それを解放することこそが求められているのです。
ソフトウェアの最適化によって、GPU がパッケージに書かれている以上の性能を発揮できる可能性を探る動きは Kog だけではありません。フランス発の ZML も、Nvidia の CUDA を迂回するハードウェア非依存のソフトウェアをリリースし、競合チップ間での高速推論を実現しました。しかしデラローは、Kog は ZML とは異なり、スタンフォード大学の研究室 Hazy Research に近い存在であり、GPU 加速に対するより深いレベルでの焦点を当てていると述べています。
デラロー自身は研究者ではなく、彼の最初のスタートアップである TechCrunch50 の 2009 年 alum Stribe は、新しい事業とは無関係です。ただし、元共同創業者で現在はベンチャーキャピタリストとなったカメル・ゼロワル(Kog のシードラウンドを主導した Varsity VC の創設者)を除けば、です。このスタートアップが深いレベルでの焦点を持つ背景には、デラローのユニークな経歴があります。
フランスのÉcole Polytechniqueで固体物理学を学んだ後、彼は攻撃的なサイバーセキュリティ(ホワイトハッカーとも呼ばれる)の世界へと進みました。Delalleau氏によれば、この経験が現在のチームに浸透させようとしている思考法を形作りました。科学分野においては、「物理法則やGPUの動作原理を理解し、それらを最大限に活用する」という考え方が根付いています。
ハッキングについては、DEFCON CTF大会で4回もファイナリストに選ばれた経験が、「アセンブリ言語やバイナリコードという極めて低いレベルまで逆解析を行い、仕組みを深く理解した上で、本来の設計意図とは異なる目的のためにそれを利用する」というスキルを教えてくれました。
このアプローチには課題もあります。非常に手作業が多く、時間がかかる点です。「新しいGPUが出るたびに、数週間から数ヶ月をかけて詳細に掘り下げ、そのハードウェアに対するGPUエンジニアリングの研究を行う必要があります。」11人のチーム体制では、これによりKogが扱えるチップの数が制限され、少なくとも当面の間はそれが続きます。
将来的には、Kog はその手法をエージェントベースのパイプラインに組み込み、より多くのチップやモデルのサポートを実現する計画です。欧州がこれらの分野で独自の能力構築を進める中、この動きは同スタートアップにとって主権を強化する追い風となる可能性があります。すでに Scaleway の支援を受けており、フランス政府系投資機関 Bpifrance や「French Tech 2030」プログラムからの後押しも受けています。
現時点では、Kog はそのアプローチが大規模言語モデル(LLM)でも機能することを世界に証明する必要があります。これがさらなる資金調達の鍵にもなります。「最初の主要モデルを 10 倍の速度で実装できれば(9 月頃になる見込み)、顧客からの反響を示し、そこからシリーズ A ラウンドへの資金調達へと進めるでしょう」と Delalleau は語っています。
*当記事内のリンクを通じてご購入いただいた場合、小規模な手数料をいただいております。これは編集の独立性には影響しません。
アンナ・ハイムはライター兼編集コンサルタントです。
アンナへの連絡や outreach の確認は、annatechcrunch [at] gmail.com までメールでご連絡ください。
2021 年以来フリーランスの記者として TechCrunch で活動しており、AI、フィンテック&インシュアテック、SaaS と価格設定、そして世界のベンチャーキャピタル動向など、スタートアップ関連の幅広いトピックを取材してきました。
2025 年 5 月現在、彼女の TechCrunch での取材は欧州の注目スタートアップに焦点を当てています。
Anna は、TechCrunch Disrupt、4YFN、South Summit、TNW Conference、VivaTech など、業界規模を問わず多様なイベントでパネルモデレーションやステージ上インタビューを担当してきました。
元 Next Web のラテンアメリカ・メディア編集者であり、スタートアップ創業者、パリ・サイエンス・ポ校出身の彼女は、フランス語、英語、スペイン語、ブラジルポルトガル語など複数の言語を流暢に操ります。
原文を表示
The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.
The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and Nvidia H200 GPUs it used for its demo.
Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch.
Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.
Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.
The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”
This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open sourced Laneformer 2B.
Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.
Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper-level focus on GPU acceleration.
Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, has nothing to do with his new one — other than his former co-founder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.
Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.”
As for hacking, the four-time finalist at DEFCON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.”
The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future.
In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program.
For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said.
*When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.*
Anna Heim is a writer and editorial consultant.
You can contact or verify outreach from Anna by emailing annatechcrunch [at] gmail.com.
As a freelance reporter at TechCrunch since 2021, she has covered a large range of startup-related topics including AI, fintech & insurtech, SaaS & pricing, and global venture capital trends.
As of May 2025, her reporting for TechCrunch focuses on Europe’s most interesting startup stories.
Anna has moderated panels and conducted onstage interviews at industry events of all sizes, including major tech conferences such as TechCrunch Disrupt, 4YFN, South Summit, TNW Conference, VivaTech, and many more.
A former LATAM & Media Editor at The Next Web, startup founder and Sciences Po Paris alum, she’s fluent in multiple languages, including French, English, Spanish and Brazilian Portuguese.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み