現代の AI 向け OSS Valkey のアーキテクチャパターンを ms から µs へ
本文の状態
日本語全文を表示中
詳細モードで約49分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
InfoQ AI/ML
NASA のスペースシャトルからの教訓を引用し、プロキシアーキテクチャが CPU コストの上昇、尾部遅延(tail latencies)の悪化、および広範囲な障害リスク(blast-radius risks)をもたらすことを指摘した。
AI深層分析を開く2026年8月6日 19:15
AI深層分析
キーポイント
プロキシアーキテクチャの隠れたリスク
NASA のスペースシャトルからの教訓を引用し、プロキシアーキテクチャが CPU コストの上昇、尾部遅延(tail latencies)の悪化、および広範囲な障害リスク(blast-radius risks)をもたらすことを指摘した。
直接アクセス Valkey の利点
直接アクセス型の Valkey アーキテクチャがマイクロ秒単位の遅延を実現し、システムの耐障害性を向上させながらインフラコストを大幅に削減できることを実証した。
AI フィチャーストアの最適化
低遅延ワークロード、特に AI フィチャーストア向けのデータレイヤー最適化において、アーキテクチャ選択がシステム全体の性能とコストに直結することを強調した。
重要な引用
proxy architectures introduce hidden CPU costs, elevated tail latencies, and blast-radius risks
direct-access Valkey architectures achieve microsecond latency, improve resilience, and slash infrastructure costs
編集コメントを表示
編集コメント
NASA のスペースシャトルの事例を技術的教訓として引き出す視点は、複雑なシステム設計におけるリスク管理の重要性を浮き彫りにしている。Valkey を用いた直接アクセスアーキテクチャは、AI 分野で急速に高まる低遅延要件に対する実用的な解決策として注目されるべきである。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
https://www.infoq.com/ から
プレゼンテーション一覧へ
From ms to µs: OSS Valkey Architecture Patterns for Modern AI
プレゼンテーションを見る
所要時間:50 分 54 秒
ダウンロード
/presentations/valkey-architecture-patterns/en/slides/Dum-1785843822385.jpg)
概要
Dumanshu Goyal は、AI フィーチャーストアのような低レイテンシワークロード向けにデータ層を最適化する方法について解説します。NASA のスペースシャトルの事例から得た教訓として、プロキシアーキテクチャが隠れた CPU コストや、尾側(テール)のレイテンシ増大、そして障害時の影響範囲拡大というリスクをもたらすことを説明。一方、直接アクセス型の Valkey アーキテクチャではマイクロ秒単位の低遅延を実現し、耐障害性を高めつつインフラコストを劇的に削減できることを実証します。
登壇者紹介
Dumanshu Goyal は AirBnb でオンラインデータの優先事項を統括しています。それ以前は Google Cloud Databases のメモリ内キャッシュを担当し、Google Cloud Memorystore のスケールと価格性能比を 10 倍改善しました。「10 倍」という数字がスライド上の約束ではなく、実際に達成された稀な事例の一つです。さらにその前には AWS で 10 年間勤務し、AWS Timestream の創設エンジニアとして活躍しました。
コンファレンスについて
ソフトウェアは世界を変えています。QCon San Francisco は、開発者コミュニティにおける知識とイノベーションの普及を促進することで、ソフトウェア開発を強化します。実践者が主導するこのカンファレンスは、チーム内でイノベーションに影響を与える技術リーダー、アーキテクト、エンジニアリングディレクター、プロジェクトマネージャーを対象に設計されています。
INFOQ EVENTS
2026 年 8 月 6 日午後 1 時(EDT)
高リスクなインシデント対応のための AI エージェント評価の構築
登壇者:ブライアン・ブジュノフスキー(Datadog AI シニアプロダクトマーケティングマネージャー)、ベンジャミン・バートン(Datadog シニアソフトウェアエンジニア)
2026 年 8 月 27 日午後 1 時(EDT)
フレームワークの奥へ:なぜエージェントの文脈はインフラの問題なのか
登壇者:ボイド・ストー(Tacnode 創設ソリューションアーキテクト)
トランスクリプト
ドゥマンシュ・ゴヤル氏:まず、NASAのスペースシャトル計画に関する物語から始めたいと思います。ここでお見せするのは、スペースシャトルが誕生する前の原初的なビジョンです。このプログラムは1970年代に始まり、2011年の運用終了まで続きました。これはあくまで机上の空論としての設計思想でした。核となるアイデアは「再利用可能な宇宙船」を造り、安価に往復飛行を実現することです。
再利用性を追求する過程で、既存の航空機と同様の要件が生まれました。つまり、「滑走路への着陸」が可能であるという条件です。この再利用性と滑走路着陸の要件から、現在のデザインが導き出されました。この設計において重要なのは、既存の航空機の構造の影響を強く受けた結果、デルタ翼を採用した点にあります。
ご覧いただいているのは、一般的に「大型デルタ翼」と呼ばれる構造です。これは実物大の模型で、「アトランティス号」という名前がついています。この宇宙船は約25年間にわたって33回のミッションを成功させ、その後退役しました。
アトランティス号の特徴であるデルタ翼をご覧いただくとわかりますが、この宇宙船が大気圏に再突入する際、激しい熱が発生します。その際、翼はこの高温の空気を切り裂くように通過します。しかし、この極度の熱にさらされるため、翼には保護が必要でした。
そこで開発されたのが、このタイルです。これらは「シリカタイル」と呼ばれる耐熱材で、写真に見えるレンガ状の小さな構造物がそれにあたります。この模型には約24,000枚のシリカタイルが貼られています。設計の目的は、宇宙船の機体全体、特に翼や腹部を熱から守り、安全に飛行させることにありました。
この航空機(宇宙船)の設計・建造当時、往復1回あたりのコストは約1,000万ドルと見積もられていました。
当初の目標は、ある程度のメンテナンスが必要ではあっても、約2週間で準備を整えることでした。しかし、実際に運用が始まってみると、予算は1,000万ドルから15億ドルへと膨れ上がり、コストが爆発的に増加しました。次のミッションに備えるまでのターンアラウンドタイムも、2週間から2ヶ月に延びてしまいました。
何が起きたのでしょうか?シリカタイルの交換が必要となりました。飛行中に予期せぬ損傷が発生したのです。「大丈夫だろう」と考えていたものの、実際には耐えきれませんでした。熱があまりにも過酷だったためです。多くの新たな発見もありました。
この問題は、地上での複雑さをさらに増幅させました。ご覧いただける通り、次のミッションに備えるために準備を整える際、地上運用は極めて複雑化しました。その複雑さは、再利用可能な時期におけるコスト面やパフォーマンス面の両方で、膨れ上がっていったのです。
2011 年、NASA がこのプログラムを退役させた際、新たな「商業乗員輸送プログラム」が始まりました。これは民間化の動きの一つです。ボーイング社の「スターライナー」や、スペースX 社の「ドラゴン」といった新型機がその代表例で、これらはカプセル型宇宙船と呼ばれるデザインです。
彼らは再利用性という根本的な要件へと回帰しました。乗員と貨物の安全を確保しつつ、従来のランディング(滑走路着陸)や耐熱タイル、翼といった複雑な要素は不要だと判断したのです。その結果、よりシンプルなカプセル型が誕生しました。
重要なのは、カプセルの先端に設けたヒートシールド(耐熱盾)を保護することです。大気圏再突入時に生じる高温から守るため、この耐熱部はカプセル本体から離して配置されています。これにより、高温の影響が本体に及ぶのを防ぎます。
この「鈍体」のデザインは成功し、その実績を積み重ねてきました。こうして宇宙船の設計思想は再構築され、「再利用可能な宇宙旅行」を実現するための新たなビジョンへと転換したのです。
この話は、ある名言を軸に展開します。これは完璧さについて語られた言葉で、アンリ・ド・サン=テグジュペリの古くからの格言です。彼はソフトウェア開発者ではなく、飛行士でありエンジニアでした。
彼が言及したのは、宇宙船のミッションにおける考え方の違いです。「ランディング用の滑走路を解決し、タイルをシステムに追加する」というアプローチと、「一歩引いて全体像を見渡し、設計そのものを問い直し、要件自体に挑戦する」アプローチの違いです。本当に必要なものなのか?そして翼やタイルを取り除き、最終結果を単純化すること。
これが私が「効率性を意識した設計」と呼ぶものです。今回の講演では、要件に戻り、表面的な理解に留まらず厳密な分析を行うことの意味について、いくつかの事例を通じてお話しします。
さらに、これらを包括的なトレードオフ分析と組み合わせることも重要です。宇宙船のミッションにおいて、この包括的トレードオフ分析は後から振り返れば明らかだったかもしれませんが、当時はタイルが実際に与える影響や、そのコストがどれほど大きくなるかという視点が欠けていたと言えます。
厳密な要件分析と包括的なトレードオフ分析を組み合わせることで、コスト効率やパフォーマンスの観点から、効率的な設計が生まれるのです。
プロフィール
私はドゥマンシュと申します。エアビーアンドビーのリードエンジニアを務め、同社のデータプラットフォームを統括しています。このプラットフォームは、エアビーで発生するすべての予約データを管理しており、その蓄積されたデータが年間の110億ドルという収益を生み出す原動力となっています。
エアビー入社前は約10年間、システムの耐久性に没頭し、いくつかの基盤システムの中身づくりに携わりました。その一つがAWS DynamoDBです。その後、私は逆に極端な方向へ進み、Googleではメモリーキャッシュによるサブミリ秒単位の超高速パフォーマンスを追求しました。
今回の登壇タイトル「msからµsへ:現代AIのためのOSS Valkeyアーキテクチャパターン」は、AIの世界における新しい現実を示しています。ここでは「性能=効率」という等式が成り立ちます。ミリ秒からマイクロ秒へと移行する際、それはリアルタイム性の確保を意味することもあれば、コスト削減が主目的となることもあります。
本講演では、RedisやValkeyといったキャッシュシステムを中心に解説しますが、Memcachedなど他のキャッシュシステムでも同様の原則が適用されるため、非常に汎用性が高い内容となっています。
ロードマップ
本日のアジェンダは以下の3部構成です。
まず第1部では、マイクロ秒(μs)レベルの低遅延がなぜ必要とされるのかについて解説します。業界全体で「より高速に」「マイクロ秒を実現したい」という声が多く聞かれますが、その背景にある真の必要性とは何かを明らかにします。AI分野での具体的なユースケースを取り上げ、マイクロ秒単位の遅延が実際に重要となる場合と、そうでない場合の違いを理解していただきます。
第2部では、ミリ秒(ms)レベルのアーキテクチャをどう運用してきたかという私たちの道のりについてお話しします。Redis と Valkey の進化を軸に、現在のアーキテクチャに至るまでの変遷を追います。さらに、アーキテクチャ選択に伴うトレードオフについても議論します。具体的には、遅延やパフォーマンスへの影響、コスト(ドル換算での支出)、そして信頼性といった観点を比較検討します。
最終的な第3部では、ミリ秒アーキテクチャに焦点を当てます。ここでは、価格対性能のバランスがどうなっているか、また信頼性の確保にはどのような要素が関わるかを解説します。
第1部:必要性の確立。まずは要件から考えましょう。なぜマイクロ秒単位の遅延が重要なのか、その理由を探ります。
今日は「AI 特徴量ストア」を例に挙げます。特徴量ストアとは何かについて解説します。この事例は DoorDash の AI プラットフォームから引用したもので、同社の公開ブログ記事に基づいています。私の講演内容と非常に親和性が高い事例です。
この事例の核心は、予測サービスにあります。ここでは 100 ミリ秒という時間予算が与えられています。マイクロ秒ではなく、ミリ秒単位です。一見すると十分な時間のように思えますが、複雑なデータ要件を考慮すると、なぜ基盤システムでマイクロ秒単位の遅延を実現する必要があるのかが理解できます。
これは DoorDash のブログ記事から抜粋したコードです。不正検出やレストランの推薦を行う予測サービスに関する記述です。同サービスにおける最優先事項は、彼らの言葉で言えば「100 ミリ秒以内に予測を返すこと」です。
この予測サービスは具体的に何を行うのでしょうか?それは「特徴量(features)」と呼ばれるものを利用します。特徴量は、リアルタイムデータを一点ずつ切り出したものだとイメージしてください。
例えば不正検知サービスのユースケースでは、「そのクレジットカードは新規か」「直近でログインは何回あったか」といった情報を確認することになります。このようなデータを実際の予測に活用し、AI モデルにリアルタイムの情報を供給して判断を下させる必要があります。
ここで重要なのは、単一の予測を行うためにシステムが数百もの特徴量を必要とする点です。多様なデータセットから多数の特徴量が集まり、初めて一つの意思決定が可能になるのです。
根本的な課題は、これをいかに高速で実現するかです。数百もの特徴量を、100 ミリ秒という時間枠内で処理する必要があるのです。これが典型的なアーキテクチャの姿となります。
本発表は、予測サービスを中心としたAI推論ユースケースについて取り上げます。アーキテクチャの詳細自体は重要ではありませんが、特に「予測サービス」に焦点を当てて議論したいと考えています。
ユーザーはこのサービスを通じてアプリと対話を行い、ここで100ミリ秒という制約が問題となります。この予測サービスはAIモデルと連携して動作しますが、同時にAI特徴ストアとも接続されています。本稿では、まさにこのAI特徴ストアの契約条件について掘り下げていきます。
具体的には、予測サービスとAI特徴ストアの間でどのような契約関係があるのか、ミリ秒単位の遅延に対応できるのか、あるいは2ミリ秒で十分なのか、それとも数百マイクロ秒レベルが求められるのかなどを検討します。
DoorDash のブログ記事に戻ると、単一の予測には数百もの特徴量が必要となるため、データを並列取得するのが自然なアプローチだと指摘されています。実際には 100 以上の異なるデータソースが存在する場合、それらの呼び出しをすべて並行して実行(ファンアウト)できます。
このように呼び出しを並行化すると問題になるのが、p99 レイテンシです。私はこれを「テールレイテンシ」と呼んでいます。この値を低く保つことが重要です。なぜなら、数百回の呼び出しを行う場合、全体の遅延は最も遅い呼び出しによって決定されるからです。
多くの呼び出しが行われる状況では、テールレイテンシはすべての予測リクエストに必ず現れることになります。定常状態での平均レイテンシが 1 ミリ秒程度だったとしても、その中で 10 ミリ秒かかるようなリクエストが一つは発生するのです。
リレーレースの比喩で説明しましょう。全員がゴールに到達し、最後の一人が着いた時点でレースが終わる場合、その遅れた一人が全体のタイムを決定します。これがさらに悪化するのが、AI ユースケースにおける複雑なデータ要件です。
この複雑さとは具体的にどのようなものか。100 回の並列呼び出しが必要で、その次に順次ルックアップを行うというパターンがあります。これは非常にシンプルですが、「並列処理」と「順次処理」が混在しています。順次ルックアップが必要な場合、データ間に依存関係があることを意味します。特定のデータを取得するには、経路をたどる必要があるのです。
注文関連の例で言えば、まずある情報を取得し、その後にのみ次の呼び出しを行えるという制約があります。単純に並列呼び出しを行っても効果はありません。順次処理が必要になる場面では、テールレイテンシ(遅延)の影響を受けやすくなります。特に高いテールレイテンシが発生すると、予算を食い潰す原因となります。
100 ミリ秒という限られた時間の中で、AI モデルが本来行うべき処理にできるだけ多くの時間を割きたいものです。しかし、単なるデータ取得のためにすでに多くの予算を消費してしまっているのが現実です。
DoorDash の AI エンジニアたちの最終的な結論をまとめましょう。彼らが説明しているのは、AI モデルのレイテンシが数ミリ秒レベルにあるという事実です。これ自体は問題視していません。当然のこととして受け入れられています。
重要なのは、その背後にあるストレージや特徴量ストア(feature store)が、モデルに提供される特徴量のレイテンシを、それに見合ったより低いレベルで処理できる必要があるということです。ここからがさらに厳しくなる部分です。例えば、モデルへの特徴量を供給する際にマイクロ秒単位での応答性が求められます。
この例では、データソースをバックエンドとして持つ特徴量ストアを Redis で構築しています。今回の発表ではまず Redis について少し触れ、その後 Valkey について詳しく解説していきます。ここで確立されたのは、特徴量ストアの要件です。前述の通り、永続的なデータソースをバックエンドに持ちつつ、マイクロ秒単位のレイテンシで特徴量を供給できることが求められています。
それでは、Redis と Valkey の簡単な紹介から始めましょう。
LLM に任せてしまう前に、まずは Redis について整理しておきましょう。平易な言葉で言えば、Redis は多様なユースケースに対応できるデータベースです。高度なデータ構造や豊富な操作機能を持ちながら、使いやすさが人気を博している理由でもあります。
この技術の歴史は古く、2009 年から始まっています。以来、その人気が着実に高まってきました。そして 2024 年、Redis のオープンソースライセンス変更という大きな転換点を迎え、Valkey が誕生しました。Valkey は事実上 Redis のフォークであり、Redis でできることはすべて Valkey でも可能です。そのため、私は両者を区別せず、文脈に応じて使い分けています。
Valkey は完全なオープンソースプロジェクトとして、コミュニティと主要なハイパースケール企業によって支えられています。これが Valkey に関する簡単な歴史的背景です。
これは Valkey のフォーク発表直後のプレスリリースからのスクリーンショットで、Linux Foundation がバックアップしています。時期は 2024 年初頭です。
また、Stack Overflow の 2025 年データベース調査結果を見ると、開発者たちが最も「欲しい」とする技術ランキングにおいて、Valkey は Postgres に非常に近い位置につけています。Postgres はすでにコミュニティでの存在感を確立していますが、Valkey もフォークからわずか 1 年でこれだけの人気を獲得したことになります。
基本紹介を終える前に、典型的な「Hello World」Valkey アプリケーションの例を見てみましょう。どのような構成になるでしょうか?クライアントと Valkey サーバーがあり、ここでは単純なキー・バリュー型のデータ取得が行われます。
これは私がターミナルで実行している Valkey CLI の様子です。CLI モードを利用できますが、アクセス方法は様々あり、一般的なクライアントツールも多数存在します。具体的な操作としては、SET my_key my_value コマンドを実行し、その後 GET などで値を取得するだけです。このように、基本となるキー・バリュー型の検索処理は非常にシンプルです。
Part 2: The Milliseconds Journey - The Architectural Choice and Its Hidden Costs
前編に引き続き、今回は DoorDash の AI ユースケースで重視した要件についてお話しします。これはミリ秒や数十ミリ秒ではなく、マイクロ秒単位で計測されるリアルタイムパフォーマンスが求められるものです。
次に、Redis における「ミリ秒の旅」について少し掘り下げていきます。アーキテクチャがどのように進化し、現在どのような状態にあるのか。そして、それが価格、パフォーマンス、レイテンシ、信頼性の観点において何を意味するのかを解説します。
こちらは DB-Engines のデータに基づくグラフです。同サイトは独自のスコアリング方式でデータベースの普及度を算出しており、過去 10 年以上にわたる Redis の人気推移を示しています。2009 年、Redis が注目を集め始めた頃まで遡ると、多くのアプリケーションが採用し始めました。当時は Valkey 101 で紹介したようなシンプルなキャッシュ構成が主流でした。つまり、アプリケーションから Redis にアクセスするだけで済む、極めて基本的な利用形態だったのです。
しかし、この人気により、アプリケーションはスケーラビリティの制約に直面し始めました。単一の Redis ノードではもはやワークロードを処理できなくなったのです。クライアント接続が多すぎたり、1 秒あたりの操作数がノードの垂直スケール限界を超えたりする事態が頻発しました。
当時、ネイティブな解決策は存在しませんでした。そこで登場したのが、私が「プロキシアーキテクチャ」や「ゲートウェイアーキテクチャ」と呼ぶ新しい設計です。これは独立した Redis ノードを連携させ、あたかも巨大な単一システム、あるいは巨大なクラスタであるかのように見せるものです。データはノード間でシャード化され、アプリケーションからはそれらに接続するだけで済みます。
私が示す青写真では、レガシーなアプリケーションの背後にゲートウェイ層を置き、さらにその下に Redis サーバーがバックエンドとして配置されています。
2009 年の状況を振り返ると、6 年後の 2015 年に Redis Cluster がリリースされました。これにより、シャード機能は大幅に改善され、Redis 自体にネイティブなシャード機能が組み込まれるようになりました。複数のノードが連携してエンドクライアントに巨大なクラスタビューを提供する仕組みです。
しかし、プロキシ層やレガシーアプリケーションは依然として残りました。トレンドは続きましたが、当時の Redis Cluster はまだ成熟しておらず、新しくて複雑でした。クライアント側で複雑なロジックが必要となるため、誰もが正しく実装できたわけではありません。そのため、すぐに普及するほどにはなりませんでした。
この状況が、プロキシアーキテクチャの採用をさらに加速させる結果となりました。レガシーアプリケーションはクラスタ機能を採用したいと考えていたとしても、クライアントの変更を避けたいという理由から、結局は中間にプロキシ層を追加し続けることになりました。
クラスターをプロキシの背後に配置すれば、システムを機能させることができます。裏側で個別ノードを独立して管理する必要はありません。クラスターを活用しつつ、プロキシがエンドクライアントアプリケーションからその複雑さを抽象化してくれます。
この成長を後押しした要因の一つは、レガシーな制約です。もう一つの理由はシンプルさや、クライアントアプリケーションとサーバーの分離(デカップリング)にあります。つまり、これがプロキシが注目されるようになった非機能的な理由の一つでもあります。
三つ目の理由はより技術的な側面です。これはレガシーからの制約というよりは、接続の多重化(コネクション・マルチプレクシング)に関わる問題です。アプリケーションのスケーリング規模が非常に大きくなり、10 万個ものクライアントポッドを稼働させるようなケースでは、裏側のクラスターへ接続を試みる必要があります。そのような大規模スケールで動作するアプリケーションは、私が知る限りでも極めて稀です。
もしクライアント側からそれぞれ 1 つずつ、合計 10 万もの接続を Redis や Valkey クラスターに直接行えば、クラスターは確実に機能停止します。そのような場合、プロキシを活用してアーキテクチャを見直す必要があります。プロキシは接続管理の負荷を引き受ける役割を果たします。
具体的には、まずクライアント側の 10 万もの接続をプロキシに集約します。プロキシはステートレスな設計のため、水平方向へのスケーリングが容易です。その後、プロキシが多重化(マルチプレクシング)技術を用いて、バックエンドサーバーとの接続数を大幅に減らしながらリクエストを処理します。
このアプローチの核心は、接続管理の責任をクラスターから切り離し、プロキシ側に移すことにあります。これが、プロキシが実際に効果を発揮してきた、あるいは現在もその役割を果たしている主要な要件の一つです。
このアーキテクチャでは、業界で広く採用されているオープンソースの Envoy プロキシを採用しています。通常 5 から 10 台のプロキシが稼働していますが、今回はそのうちの一台を例に挙げています。ただし、ここで取り上げる内容は Envoy に固有のものではなく、議論全体を通じてこの例を引き続き使用します。
この例では、Envoy プロキシ層がロジックを実行するホストとなります。このロジックは主に 2 つの部分から構成されます。1 つ目は、裏側で独立した Redis または Valkey ノードが存在するパターンです。今回は Redis から Valkey に切り替えています。Envoy は単純なハッシュスキームを用いてクライアントデータをシャード化できます。データベースではないため、わずかなデータ消失の懸念は残りますが、実用には十分機能します。
2 つ目はクラスタモードです。このモードでは Valkey 自身がデータのシャード構成を把握しており、Envoy プロキシもこれに接続可能です。これらのプロキシの最大の利点は、背後にあるバックエンドシステムと複数のプロトコルをやり取りできる点にあります。
Redis の歩み、プロキシの登場とその存続理由、そしてプロキシを介したアーキテクチャが採用される主な背景についてお話ししました。ここでは、価格、パフォーマンス、レイテンシといった現実的なトレードオフや指標に焦点を当てていきます。実際の様子を見ていきましょう。
まずはレイテンシから始めます。これはパフォーマンスの観点で捉えることができます。今回の検証では、Envoy ノード 3 台と Valkey ノード 3 台からなるアーキテクチャを対象にしますが、分析をわかりやすくするために構成を簡略化します。具体的には Envoy を 1 台、Valkey も 1 台に減らしてシミュレーションを行います。クライアントがプロキシに接続し、裏側で Valkey ノードと通信するシンプルな構成です。
今回のベンチマークでは、memtier_benchmark というユーティリティを使用しています。
これは特別な機能があるわけではなく、負荷をかけるためのツールで、システムの特性を測定するものです。この環境は AWS の EC2 にデプロイされており、使用している VM タイプを確認できます。プロキシと Valkey はどちらも全く同じマシンタイプを使用しており、同じアベイラビリティゾーン内に配置されています。クロスゾーン間の通信を避け、干渉を最小限に抑えるためです。両方とも 8 コアの構成です。
これは私のターミナルから実行したベンチマークの結果です。グレーアウトしている詳細な情報にはこだわらず、重要なのは memtier_benchmark を使用している点と、実務で一般的に見られるような読み取り中心のワークロードであることです。このケースでは、読み書きの比率はおよそ 90:10 です。トラフィックの 90% 以上が読み取り操作です。
この構成では、半分の QPS(秒間クエリ数)を生成しています。具体的には、秒間 50 万クエリを処理しており、そのうち 90% 以上が読み取りで、残りがセット操作です。
次にレイテンシについて説明します。この秒間 50 万クエリの状況における遅延時間を測定しました。2 つの指標があります。1 つはパイク(p99)と呼ばれる尾部の遅延で、これはおよそ 2.5 ミリ秒です。これは 1〜10 ミリ秒程度の範囲に属するミリ秒帯域の値です。もう 1 つが p50、つまり中央値の遅延時間で、これは約 1 ミリ秒です。
これを「ミリ秒アプリケーション」と呼んでいます。これらの数値は非常に優秀です。秒間 50 万クエリを処理しながら、レイテンシは極めて低く、単一桁のミリ秒で抑えられています。決して悪くありません。
これらのセットアップそれぞれについて、レイテンシ、価格、パフォーマンスの側面を一つずつ詳しく見ていきましょう。まずは p50 レイテンシに焦点を当てます。
システムが現在どのような状態にあるのかを理解するためには、その特性を把握することが重要です。p50 レイテンシについては 1 ミリ秒という数値をお話ししましたが、この値はクライアント側から測定されたものです。つまり、1 ミリ秒で完結する完全な往復通信の時間です。
このプロセスにはネットワークホップが 2 回含まれています。各ホップのコストを測るため、Linux ユーティリティである My Traceroute を使用しました。これは ping に似たツールですが、より高度なオプションを備えています。
結果を見ると、クライアントとプロキシノード間の 1 つのホップに 300 マイクロ秒かかります。この 300 マイクロ秒の区間は、基本的に同じアベイラビリティゾーン内にあります。
先ほどお話しした通り、干渉はほとんどなく、クロスゾーンホップによる遅延も発生しません。Envoy から Valkey までの往復で、同じく 300 マイクロ秒がかかります。つまり、合計 600 マイクロ秒(片道 300 マイクロ秒×2)がネットワークに消費され、p50 レイテンシの 1 ミリ秒のうち半分近くを占めています。
残りの 400 マイクロ秒には、リクエストをルーティングするためのプロキシ処理全体と、データセットを提供する Valkey サーバーの処理が含まれます。さらに、クライアント側でリクエストを送信し、レスポンスを受信して所要時間を計測する処理も含まれています。
これらの一連の処理は、すべて 400 マイクロ秒以内で完了します。リモート参照におけるネットワーク遅延を完全にゼロにすることは難しいですが、これが p50 レイテンシにおける遅延の内訳です。
次に、特定の処理能力(クエリ数/秒)を達成するために必要なコストについて見ていきましょう。コストは非常に重要です。エンジニアとしては通常、レイテンシやパフォーマンスについて議論しがちですが、私は「ドル」についても語ることを好みます。今回はその点に焦点を当てます。
先ほどの設定では、2.5 ミリ秒の尾部レイテンシで 50 万 QPS を達成しました。これは非常に優れた結果です。この実行中の CPU 使用率を見ると、Envoy で 90% に達しており、これはかなり高い数値です。つまり、ほぼ限界まで使い切っていたことを意味します。一方、Valkey サーバーの CPU 使用率は約 60% でした。まだ多くの容量が余っており、さらに活用できる余地があります。
先ほどの図を思い返すと、視覚的にも Envoy が 90%、Valkey が 60% であることがわかります。なぜプロキシ(Envoy)だけがこれほど高い負荷がかかるのでしょうか?なぜこれほど計算集約的な処理となっているのでしょうか?一体どれだけの CPU を消費しているのでしょうか?
キャッシュの世界では、データベースの世界とは異なる点が見えてきます。データベースでは QPS(1 秒間のクエリ数)よりも、レイテンシや I/O といった要素を重視する傾向があります。
一方、キャッシュの世界で重要視されるのは主に 2 つの要素です。1 つ目は QPS です。アプリケーション側がデータベースへの負荷をキャッシュにオフロードする場合、キャッシュは高い QPS を処理するか、あるいは大量のメモリストレージを必要とします。メモリストレージは一般的なディスクベースのストレージに比べてコストが高いため、キャッシュワークロードでは QPS の側面にも注意を払う必要があります。
この高 QPS によるトラフィックはプロキシを経由して転送されます。ここでは、計算リソースがどこで消費されているのか、その全体像を理解することが重要です。具体的には、Envoy における 90% のリソース消費の内訳について把握しておくことが必要です。
遅延について議論する際、ネットワークホップがその大半を占めることが分かりました。CPU 側の観点から言えば、プロキシを経由してトラフィックを転送すると、基盤となる Valkey サーバーに比べて入出力(I/O)の量が倍になるという直感があります。
Valkey サーバー自体を見れば、リクエストを受け取ってレスポンスを送り返すだけです。しかしプロキシの場合は少し複雑になります。まずクライアントからリクエストが来ます(1 回目)。次にそのリクエストが Valkey サーバーへ送られます(2 回目)。Valkey サーバーからのレスポンスが戻ってきます(3 回目)。最後にそのレスポンスがクライアントへ返されます。
つまり、プロキシノードは、通常の Valkey の I/O で想定される量の倍のデータを移動、あるいはネットワークパケットを転送する必要があります。このネットワークパケットを移動させる一連のプロセスは、非常に計算集約的な処理です。ここで示されているのは、リクエストをルーティングするプロセスの流れです。
その下層では、ネットワークカードがパケットを受け取ると、まずホストメモリへ格納しなければなりません。これは仮想マシン上で動作しているため、物理ホスト上に存在するからです。一旦物理マシンのメモリに入ると、別のプロセスがそれをゲスト VM、つまりゲスト OS へ転送します。その後、割り込みや何らかのメカニズムを通じてカーネルの TCP スタックへと届けられ、最終的にリクエストルーティングコードを実行しているアプリケーションに到達します。
データを送信する際は、この経路を逆順でたどることになります。
これは直感(インサイト)を構築するための話でした。I/O 処理が倍増したという点については理解いただけたと思います。次は、実際のデータを見てみましょう。ここでプロファイラーを用いて、先ほどの直感を検証します。
フラムグラフも私が最も好きなツールの一つですが、スライド資料として提示するには読み解くのが非常に困難です。フラムグラフは、実行中にスタックフレームをキャプチャし、他のフレームと比較して頻繁に出現するものを特定する可視化ツールです。特定の領域が他よりも頻繁に現れる場合、十分なサンプル数を取得することで、その部分が他の部分よりも多くの CPU 時間を消費している可能性があると判断できます。
ただし、スタックフレームの詳細な読み込みについてはご安心ください。ここではより高レベルな分析を行います。90% の Envoy CPU 消費率について確認しましたが、プロセス側でも計測可能な内訳があります。それは「ユーザーモード」と「システムモード」です。ユーザーモードとは、アプリケーションが実行されている領域のことです。
プロキシを実装するコードは、ユーザーモードで実行されます。その後、システムコールや TCP スタックとやり取りしますが、これらはすべてカーネルモードに属します。
この 90% の CPU 使用率の内訳は非常に均衡しており、ユーザーモードとカーネルモードがそれぞれ約 45% を占めています。アプリケーション側では I/O が頻繁に行われ、当然ながら多くの処理が発生しています。
図の大きな二つのボックスはフレイムグラフの一部です。左側の部分はアプリケーションの処理プロセスを示しており、I/O そのものというよりは CPU 負荷が大きい領域です。ここはコードを書き換えるなどして、大幅な最適化が可能です。
右側は TCP の送信・受信パスを表しています。ここではアプリケーション側の一部と、カーネルフレームが含まれています。先ほど述べた通り、I/O は他の処理に比べてコストが約 2 倍かかります。この I/O が、CPU 使用率の 90% を占める大きな部分を形成しているのです。
私たちが理解しようとしている重要な点は、「プロキシが CPU を大量消費する」と言った場合、実際には何をしているのかという点です。実は主にリクエストのルーティングを行っています。これに相当量の CPU が費やされています。残りの CPU 使用率については、Envoy プロキシの場合、その半分が I/O に使われています。
Valkey サーバーに十分な容量があるという前提で、現在の構成は 90% が Envoy に依存している状態です。つまり、Envoy がボトルネックになっていると考えられます。
そこで提案するのは、このレイヤーを水平方向にスケールさせるために、さらに 1 つの Envoy プロキシを追加することです。2 つの Envoy プロキシはステートレスなので、リクエストをロードバランシングできます。クライアントは任意のプロキシを選択してトラフィックを送信します。
その後、再度ベンチマークを実施しました。サーバー前方に 2 つのプロキシを配置した結果、システム全体で 100 万 QPS を処理できるようになりました。以前は Envoy プロキシが 1 つのとき、システムを通じて流れる QPS は 50 万でした。これが 2 つになると 100 万になります。
これは素晴らしい結果です。ハードウェアを 2 倍にしたことで処理能力も 2 倍になり、優れたスケーラビリティ特性を示しています。
レイテンシは非常に似ています。p50 は再び 1 ミリ秒、タイレイト(最悪の遅延)も 2 ミリ秒超で、これまでの結果とほぼ同じです。ミリ秒単位の応答性を維持できています。このシステム構成では、1 秒間に 100 万回のクエリ(QPS)を処理できることも確認できました。
コストについて触れたので、この構成にかかる費用を見てみましょう。私が使用している第 8 世代の Graviton インスタンスは月額約 230 ドルです。プロキシ用に用意したハードウェアは約 400 ドル、Valkey の部分は 200 ドル強です。つまり、低単位のミリ秒レベルのタイレイトで 1 秒間に 100 万回の QPS を処理するには、月額約 700 ドルが必要となります。
これは重要な結果です。一瞬立ち止まって確認しておきましょう。後ほど議論の中で再び取り上げます。決して悪い結果ではありません。アーキテクチャの詳細を省略して言えば、これはインメモリ型の設定です。インメモリは高価であることは誰もが知っています。
月額 700 ドルはビジネスにとって正当な支出に見えますが、実際にはこの金額よりも遥かに多くかかる可能性があります。具体的には容易に 1,500 ドル近くになるでしょう。なぜなら、この構成には高可用性(HA)の機能が組み込まれていないからです。高可用性を備える場合、レプリカのコピーが必要になり、ゾーン間の通信も発生します。その結果、ネットワークジッターが増え、コストもさらに嵩みます。そうなれば、レイテンシは悪化し、費用も増大するでしょう。
ここで話しているのは、それらの変数を排除した同一アベイラビリティゾーン内での、意図的に最適化された構成の話です。
これまでお話しした通り、レイテンシの観点について解説しました。また、特定のレイテンシ条件下で 100 万 QPS を処理する際のコストについても触れました。
通常、プロキシを導入した時点でアーキテクチャ評価はここで終わります。私たちはプロキシ経由のアーキテクチャを分析していたのです。このシステムはステートレスです。ステートレスなシステムは、これまでの経験上、あまり問題を起こさないものです。そのため、セットアップを超えてさらに深く検討することは稀でした。
しかし、3 つ目の要素である「信頼性」には、興味深い落とし穴が潜んでいます。それは単一障害点(Single Point of Failure)の問題です。私たちの設定にはまさにこのリスクが存在します。私が初めてこの問題に直面した際、なぜこのアーキテクチャで信頼性が損なわれるのかは、直感的にはわかりませんでした。
その原因は「ヘッド・オブ・ライン・ブロッキング」や単一障害点に関連しています。これらについて詳しく見ていきましょう。これらの落とし穴を理解するためには、実際にトラフィックを流す必要があります。特別なことはせず、基本的な負荷テストを行います。
アーキテクチャは整いました。今回は単一ノード構成を一旦手放し、プロキシ 3 台と Valkey サーバー 3 台の構成に戻します。ノード数が増えると興味深い挙動が見られるからです。
ここで少量のトラフィックを流して、クライアント側の可用性がどう見えるかを確認します。今回はベンチマーク用のスクリプトではなく、カスタムスクリプトを使います。現在流しているのはごく少量のトラフィックで、表示されている可用性は 1 秒ごとに 100% です。つまり、毎秒 100 QPS のリクエストがすべて成功しています。
しかし、失敗や最悪の場合は遅延もつきものです。そこで今回は、あるシャードに意図的に遅延を注入し、システムがどう振る舞うかを見てみましょう。
これは Lua スクリプトを使った手法です。Redis にはこの強力な機能があり、Valkey でも同様に利用できます。サーバーサイドで関数を実行し、特定のロジックを注入して動作させることができます。ここで示すコードは、メインサーバーを約 5 秒間遅延させるためのワークアラウンドです。これにより、意図的な遅延や障害を発生させます。
この実験は簡単に再現できますが、本番環境では絶対に実行しないでください。本番サーバーにとっては致命的な影響を与える可能性があります。
次に、クライアント側の可用性がどうなるかを見てみましょう。通常時は稼働率が 100% ですが、スロワーを注入した 4〜5 秒の間は可用性がゼロに低下します。この 5 秒間は、すべてのリクエストが通じず、完全に利用不能な状態になります。
今回のテスト環境ではトラフィックは均等に分散されていました。つまり、スロワーの影響を受けた shard_1 には全体の 3 分の 1 のトラフィックが割り当てられていました。一方、他のシャード(shard_2 と shard_3)は本番環境でもテスト環境でも一切変更せず、本来通り全体の 3 分の 2 のトラフィックを受け取るはずでした。しかし、結果として全体の可用性はゼロになってしまいました。
さらに注目すべき点は、すべてのトラフィックが shard_1 に集中していることです。shard_2 と shard_3 は実際には全くトラフィックを受けていません。これは「シャード自体は正常で CPU も消費していないのに、なぜトラフィックが来ないのか?」という疑問を投げかけます。
実際のサーバーを運用している状況を想像してみてください。ノードが 3 つだけではなく、プロキシとシャードの両方で数十、あるいは数百ものノードが存在し、シャード 1 つがダウンすればクラスター全体が停止してしまうようなケースです。今回はその根本原因を解明します。
クライアント側から可用性がゼロであるとの報告があったため、まずはクライアントにズームインして状況を把握しましょう。健全な状態ではどのような様子になるか見てみます。ベストプラクティスに従い、クライアントはプロキシを経由して接続する必要があります。接続数が必要となるため、コネクションプールを活用します。都度新しい接続を作成するのはオーバーヘッドが大きいため避けるべきです。このコネクションプールの上限を 5 に設定しています。
この 5 つの接続が、シャード_1 に対するリクエストやその他のリクエストを処理します。これを超えるリクエストはスロットリングされます。シャード_1 へのリクエスト応答時間が長くなると、他のシャードへのリクエストがリソース不足に陥り、実行できなくなる(スターブ)状態になります。
接続プールがすべて shard_1 のリクエストで占有されてしまうのは、これらのリクエストが遅く、接続を早く解放しなかったことが原因です。これがサービス停止を引き起こします。ご覧の通り、shard_1 へのすべてのリクエストがクライアント側の接続プールを独占してしまいました。shard_2 や shard_3 にはリクエストが届かず、すべて拒否されています。このように、単一のシャードがシステム全体を完全なダウンに陥らせてしまうのです。
これらのプロキシ型アーキテクチャから得られた教訓をまとめると、700 ドルのコストで 1 秒間に 100 万回のクエリ(QPS)を低遅延で実現できる一方、一つのシャードがクラッシュするとシステム全体が停止するリスクがあることがわかりました。
Part 3: The Microsecond Playbook - The Strategic Path to µs Efficiency and Resilience
最後のテーマはマイクロ秒の世界です。ここからはアーキテクチャの詳細に踏み込みましょう。
ワナー・ヴォゲル博士の言葉が思い出されます。「これまでやってきた方法」を安易に信じず、現在のシステムに対して疑問を持ち続けることが重要です。常に最適化と改善を追求する姿勢こそが、技術進化の鍵です。
スペースシャトルの例を考えてみましょう。ランディング用の滑走路という制約を問い直すことで翼を廃し、はるかに効率的なカプセル型宇宙船へと進化させました。私は今回も同じ発想で、プロキシという要素を一旦取り除いてみます。特定のユースケースではプロキシが不可欠な場合もあるでしょうが、議論の前提として今回はそれらを排除し、その影響を確認します。
プロキシがなくなった後、クライアントはサーバーに直接アクセスできるようになります。ただし複数のサーバーが存在する以上、「データがどこにあるか」を把握する必要があります。
必要なのは、クラスタのトポロジを理解できるスマートなクライアントです。Valkey クラスタには、クライアントがそのトポロジを学習し、リクエストを適切にルーティングするための仕組みが用意されています。
今回は、レイテンシ、コスト、パフォーマンス、信頼性の各側面について検証しますが、順序は逆に行います。まず信頼性から始め、次にコスト、最後にレイテンシについて解説します。
プロキシアーキテクチャと直接アクセスアーキテクチャの両方で検証を行うため、同じスクリプトと設定を用いて、あるシャードに意図的に遅延を注入します。
その結果、以前のような完全な停止(可用性ゼロ)ではなく、部分的な障害が発生していることが確認できます。可用性は 60% から 80% の間を推移し、これは期待通り全体の約 3 分の 2 のトラフィックが影響を受けていることを示しています。シャード_1 が故障した際、報告された可用性は 66% でした。
では、先ほどの根本原因分析と同様に、クライアント側の挙動に焦点を当てて詳しく見ていきましょう。
クライアントがシャードに直接接続する場合、エンドポイントごとに異なるコネクションプールを設けるだけで済みます。各コネクションプールには、そのシャードへリクエストを送るための多数のコネクションが含まれています。
ここで問題となるのは、shard_1 が遅延を起こした際です。すると、該当するコネクションプールのすべての接続が占有され、新たなリクエストが殺到してもスロットリング(制限)がかかります。重要なのは、shard_2 や shard_3 からのリクエストは保護され、完全に分離されている点です。
シャード間が互いに独立しているため、シャード側で実現される障害隔離(fault isolation)の概念は、クライアント側にも拡張されます。つまり、他のリクエストへの影響を最小限に抑えられるのです。
これは「バルクヘッドパターン」を思い起こさせます。このパターンは古代ギリシアの船にまで遡ります。当時の船には、区画を分ける隔壁(バルクヘッド)が設けられていました。現在でも、その形状や形式は多岐にわたって採用されています。
これらのバルクヘッドは基本的に船内の区画のようなもので、船体の構造的な完全性を保つ役割を果たしています。もう一つの重要な機能は、障害隔離です。シャード化されたアーキテクチャと同様に、一部の部分が損傷しても、それが原因で船全体が沈むわけではありません。確かに影響はありますが、船は依然として浮いたまま航行を続けることができます。
先ほどは信頼性について解説しました。プロキシアーキテクチャでは、1 つのシャードが低速になるだけでクラスター全体が停止してしまう一方、ダイレクトアクセスアーキテクチャでは期待通り障害を隔離できることを確認しました。
次に注目するのはコストとパフォーマンスです。ここでは Valkey シャード 1 つにプロキシなしで構成した環境でベンチマークを実行します。この構成で 100 万 QPS を達成可能です。Valkey サーバー自体は約 230 ドルという低コストでありながら、プロキシを介さずに 100 万 QPS の処理能力を発揮できます。
この Valkey 環境でのレイテンシ数値も以下の通りです。パケットの末尾(tail latency)では 2 ミリ秒を超えることなく、約 567 マイクロ秒を記録しました。これは 100 万 QPS という高負荷下での結果です。
まとめと参考文献
まとめると、プロキシなしで同じ 100 万 QPS を実現できるため、コストは 3 分の 1 に削減できました。一方、プロキシを介在させた構成では末尾レイテンシが 4 倍に悪化し、Valkey サーバー単体の約 600 マイクロ秒に対し、約 2.4 ミリ秒となりました。
本発表で使用した参考文献は以下の通りです。スペースシャトルの事例や、ドクター・ヴェルナー・ボーゲルの「節約アーキテクチャの法則」、そして Valkey の 10 億 RPS を実現するイノベーションについて触れました。今回は深掘りできませんでしたが、これらの文献でさらに詳しく学ぶことができます。
主な学び
最後のスライドでは、主要なポイントをまとめました。ネットワークホップによる遅延が、p50 の 360 マイクロ秒のうちほぼすべてを占めていた 300 マイクロ秒であったことを理解し、直接アクセスアーキテクチャへと移行した経緯です。
このアーキテクチャがコストに与える影響も確認できました。プロキシ型アーキテクチャでは 700 ドルだったものが、直接アクセス型ではわずか 230 ドルに抑えられています。また、プロキシ型アーキテクチャにおける単一障害点の問題も、直接アクセス型によって解消されました。
最後に、Valkey サーバーがこれほど効率的である理由は、裏方として貢献してくれたコミュニティの取り組みにあります。これらの活動が、インフラコストをさらに低下させる原動力となっています。
要約付きプレゼンテーション をもっと見る
録画日:2026 年 8 月 6 日
原文を表示
From ms to µs: OSS Valkey Architecture Patterns for Modern AI
View Presentation
Speed:
Download
50:54
/presentations/valkey-architecture-patterns/en/slides/Dum-1785843822385.jpg)
まとめ
Dumanshu Goyal discusses optimizing data layers for low-latency workloads like AI feature stores. Drawing lessons from NASA's Space Shuttle, he explains how proxy architectures introduce hidden CPU costs, elevated tail latencies, and blast-radius risks. He demonstrates how direct-access Valkey architectures achieve microsecond latency, improve resilience, and slash infrastructure costs.
Bio
Dumanshu Goyal leads Online Data Priorities at Airbnb. Previously, he led in-memory caching for Google Cloud Databases, delivering 10x improvements in scale and price-performance for Google Cloud Memorystore, one of the rare times “10x” was more than a slide promise. Before that, he spent 10 years at AWS as the founding engineer of AWS Timestream.
About the conference
Software is changing the world. QCon San Francisco empowers software development by facilitating the spread of knowledge and innovation in the developer community. A practitioner-driven conference, QCon is designed for technical team leads, architects, engineering directors, and project managers who influence innovation in their teams.
INFOQ EVENTS
- August 6th, 2026, 1 PM EDT
Building AI Agent Evals for High-Stakes Incident Response
Presented by: Brianne Bujnowski - Senior Product Marketing Manager, AI at Datadog, and Benjamin Barton - Senior Software Engineer at Datadog
- August 27th, 2026, 1 PM EDT
Below the Framework: Why Agent Context Is an Infrastructure Problem
Presented by: Boyd Stowe - Founding Solutions Architect at Tacnode
Transcript
Dumanshu Goyal: I would like to open with a story. This story is about NASA's Space Shuttle program. What you see here is the original vision behind the space shuttle. This program started back in 1970s, continued all the way to 2011 until it was retired. This is an on-paper vision. The idea was to build a reusable spacecraft. The idea was to go back and forth into space and do it cheaply. This was basically the designs behind the spacecraft. Then from the reusability, they came up with this requirement that's similar to the other planes we see out there. We want runway landing, the spacecraft to be able to land on a runway. The runway landing from reusability became a requirement, and that's where they came up with this design. What's important about this design is because of being influenced with the existing structures of planes, they added these delta wings.
What you see there is basically called a large delta wing. This is a real model. Its name is Atlantis. It did 33 space missions over about 25 years, until it was retired. This Atlantis, again, you can see the delta wings over there. What happens is when these spacecrafts, they reenter the atmosphere, there is high heat generated. Then the wing basically slices through that heat. Now, because it's experiencing that heat, it needs to be protected. This led to the requirement of installing these tiles. These are silica tiles. What you see in the picture there, the small brick-shaped structures, these are the silica tiles. This model has about 24,000 silica tiles on it. The idea was to protect the body, the wings, the underbelly of the spacecraft and make it work. When they designed and built this airplane, it was estimated to be a $10 million round trip, what it would cost us.
Then it would require some maintenance and it would be ready in about two weeks' time frame. This was the original goal. When they started using it, in reality, the $10 million ended up to be $1.5 billion. The costs exploded. The two weeks' turnaround time to get ready for the next mission got to be two months. What went wrong? The silica tiles, they had to be replaced. There were unexpected damages during the flight. They were thinking that, ok, they would hold it well, but they were not. The heat was too extreme, a lot of new discoveries. This problem ultimately added to the on-ground complexity as well. The on-ground operations, as you can see here, they got really complex when they had to get this ready for the next mission. The complexity also bloated up not just the cost aspects or the performance aspects around when it can be reused.
In 2011, when NASA retired this program, another program started, the Commercial Crew Program. This is a bit about privatization. This is where you see Starliner there from Boeing. Another top model is called Dragon from SpaceX. These are what we call as capsule designs. They went back to the core requirement of reusability. There was reusability. You want to keep the payload, the crew safe. There is safety of the crew as the core requirement. Then they basically challenged the runway landing requirement. Is that really required? Do we need to really land this on a runway? They let go of that requirement. They let go of the tiles. They let go of the wings. They came up with a simpler design, a capsule. All they had to do was to protect the nose of this capsule, which has this heat shield on the top, so that when it enters the atmosphere, it can sustain high heat.
Not just sustain it, but also move it away from the body of the capsule, so that the rest of the capsule is not impacted. This blunt body of this capsule was a success, and it has been building up on it. This is how they went back to the requirements, and then basically reshaped the spacecraft vision around how do you do reusable space travel.
I want to frame this story around a quote. This quote is about perfection. This is an age-old quote from Antoine de Saint-Exupery. He was an aviator, an engineer, not a software developer like us, or most of us. What he mentioned, when we talk about the spacecraft mission, you have this mindset, I need to solve for runway landing. I'll add the tiles to the system. Then you keep fixing it, keep maintaining it. Versus taking a step back, taking a holistic look, coming up with the capsule design, challenging your requirements. Do you truly need them? Then basically getting rid of the wings, the tiles to simplify the end result. This is basically what I refer to as designing for efficiency. In this talk, we'll cover several examples of what it means to go back to your requirements, do a rigorous requirement analysis, not take them at the surface.
Also, combine them with a holistic tradeoff analysis. This tradeoff holistic analysis, what you could say in the spacecraft mission, it's hindsight now, but you could say it was missing as to what would be the real impact on the tiles. How much it would cost us for real. When you combine this rigorous requirement analysis and the holistic tradeoff analysis, that's when you come up with the efficient designs, whether it's cost-effective or for performance reasons.
Professional Background
I'm Dumanshu. I'm a lead engineer at Airbnb. I lead their data platform. This data platform is responsible for all the bookings you do on Airbnb. All that data gets stored. This is what drives their $11 billion revenue. Before Airbnb, just a little bit about my background. I spent a decade obsessed with durability, where I got to work on internals of several foundational systems. One of them was AWS DynamoDB. After that, eventually, I moved on to the other extreme, which was raw sub-millisecond performance with in-memory caching at Google. This is my talk's title. It shows us basically the new reality about AI, where performance equals efficiency. When we are moving from milliseconds to microseconds, sometimes it's about real-time performance. We'll see how in the AI space this matters. Sometimes it's not about the latency or the performance. Most of the time, it would be about the cost. We'll also go over how moving from these milliseconds to microseconds affects your cost. I'll frame this talk around Redis and Valkey, some of the caching systems, but it's pretty generic in the sense it applies to maybe Memcached or a bunch of other caching systems you might be using.
Roadmap
This is my agenda. We'll cover it in three parts. The first part is about the need of microseconds. We all talk about, we want to get faster. We want microseconds. What's the real need behind microseconds? We'll do a case study, a use case in the AI space to understand why microsecond latency could be truly important, or it might not be in your case. Then the second part talks about our journey of how we are living with these milliseconds architectures, how it got there. I'll frame it around the evolution of Redis and Valkey. Then we'll also cover the architectural tradeoffs, like what's the impact on latency, the performance, the cost, the dollar that you pay for the architecture, and then reliability. Then, in the end, the third part will cover the milliseconds architecture. In the milliseconds architecture, we'll talk about what the price performance equation looks like there, what the reliability equation looks like there.
Part 1: The AI Data Wall - Why Milliseconds Are No Longer Good Enough
Part one, establishing the need. We want to go with requirements. We'll start with what's the need of these microsecond latencies. Why does it matter to us? The example I want to use today is an AI feature store. I'll talk about what the feature store is. This example I'm taking from the DoorDash AI platform. All public blogs. Just was quite relatable to my talk. This example basically starts with a prediction service where you have a 100-millisecond budget, not microsecond, 100-millisecond budget. That's a lot of time, to basically drive a prediction. Then when you combine that with complex data needs, you'll see why it requires an underlying system to serve microsecond latency. This is code from one of the DoorDash's blog posts. This is about a prediction service which is used to detect fraud or provide restaurant recommendations. The number one requirement of this prediction service, as they quote, is basically to be able to serve a prediction within 100 milliseconds.
What does this prediction service do? It looks as something called as features. You can think of features as a piece of real-time data. If you are in the fraud prediction service use case, you would see, is this credit card new? How many logins have happened recently? Stuff like that. You want to be able to feed real-time information to your AI model to be able to drive a prediction out of it. The important thing to note here is for a single prediction, the system needs hundreds of features to be able to make that prediction. It's like quite a diverse dataset, lot of features, these pieces of data coming together to actually make a decision out of it. The underlying challenge is, how do you make it happen so fast? You're talking about hundreds of features. You're talking about 100 milliseconds. This is like a typical architecture.
It's an AI inference use case around the prediction service. The architecture details are not that important, but I think I would like you to focus on the prediction service which basically gets this user contract. The user interacts with the prediction service through the app, and this is where the 100-millisecond constraint comes into picture. This prediction service works with the AI model. Then it also works with this AI feature store. This AI feature store is the contract that we are trying to dive into. Like, what's the contract between the prediction service and the AI feature store? Can it serve millisecond latencies? Is 2 milliseconds good enough, or does it have to be in hundreds of microseconds?
Going back to the DoorDash blog post, this is where they mention, because the single prediction needs hundreds of features, so a natural thing to do is to fetch data in parallel. I have 100 different sources or whatever. I could basically just fan out those calls. When you fan out those calls, what they call out here is the p99 latency. I basically refer to it as tail latency. You need to keep the tail latency low. Why does that matter? Here is a relay race example. When you're basically making hundreds of calls, your latency is decided by the slowest call. When you are having so many calls, your tail latency is bound to show up in every single prediction request. Your latencies, even though in steady state, might be just about 1 millisecond, but there would be that one request that maybe goes all the way to 10 milliseconds.
The analogy I want to map to the relay race where you need everyone to finish, get to the finish line, and then one person basically is lagging behind, and that dictates what happens, how soon you can finish the race. This gets worse. We talked about the AI use case. We also talked about there is a complex data need. What does this complex data need look like? We went over, you need to make 100 parallel calls. Then the next one is a sequential lookup. It's pretty simple. There is a parallel thing. There is a sequential thing. The sequential lookup, all it means is you have data dependencies. In order to get to a piece of data, sometimes you have to traverse the path. Here is an order-related example. You have to fetch a piece of information, and then only you can make the next call. Simply making calls in parallel doesn't help.
Sometimes you have to make it sequentially. When you make calls sequentially, and you are prone to these tail latencies, the high tail latencies, that's where you'll see it starts to eat away your budget. Out of that 100 milliseconds, you want to give as much time as possible to the AI model to do the real thing. Just for fetching the data, you have already eaten up a lot of budget, in this case.
This brings us to the final verdict from the DoorDash AI engineers. What they basically describe here is their AI model latency is in low milliseconds range. We are not worried about that. That's a given. What that means is the underlying store, the underlying feature store, needs to be able to serve the features at a latency which is proportionately lower. This is where it gets even tighter, where we talk about microseconds, for example, to serve the features to this model. Another code, in this example, they build this feature store backed by a data source using Redis. For our talk, we'll talk a bit about Redis. Eventually, we'll go into Valkey. What we have established here is the requirements for this feature store. Like I said, it's backed by another durable data source. Then we are talking about serving these features at a microsecond latency level. I would like to start with a quick introduction to Redis, Valkey.
I'll get back to our favorite LLM to do it for us. Redis, I think in plain, simple terms, it serves numerous use cases, you name it. It has advanced data structures, so many operations available. It's quite simple to use, and that's what basically speaks to its popularity. The important thing to note is this technology already goes back to 2009, since when it started picking up on its popularity. Then, in 2024, that's when the open-source licensing change happened with Redis, and Valkey was born. Valkey is effectively a fork of Redis. You can do all what you want to do with Redis with Valkey as well. That's why I use it interchangeably. Then, it's fully open source. It's backed by the community, a bunch of hyperscalers. That's a little bit history about Valkey. This is a screenshot from the press release around when the Valkey fork was announced, Linux Foundation backed it up. This is early 2024. This is a Stack Overflow 2025 database survey results. What you see here is, in terms of the desired level among these developers, Valkey is trending very close to Postgres. Postgres has really gained community presence, and in terms of the desired level, just this is within one year of its fork.
Before we end the basic introduction, I want to do a typical 101, like a Hello World Valkey application. What does it look like? You have a client. You have the Valkey server. It's a plain, simple key-value lookup, what we have here. This is my terminal running Valkey CLI. You can use the CLI mode. There are a bunch of ways to access it, typical clients. What you do is a SET my_key my_value, and then you basically can get it after that. It's as simple as that plain key-value lookup. We're not talking about anything more advanced than that.
Part 2: The Milliseconds Journey - The Architectural Choice and Its Hidden Costs
Going to our part two. So far, what we talked about is how the DoorDash AI use case that we looked at is basically a requirement for real-time performance, which is measured in microseconds, not milliseconds or not tens of milliseconds. Next, we are going to talk a bit about our milliseconds journey. In the context of Redis, how the architecture has been evolving, what it's like today, and what does it mean in terms of price, performance, latency, reliability. This is a chart from DB-Engines. It's widely popular. They have a way to calculate scores. This showcases, over the last 10-plus years, how the popularity of Redis has been rising. Going back to all the way to 2009, when Redis started to get popular, a lot of applications started using it. They had a simple caching setup, like what we saw with the Valkey 101. In this case, your application is, again, accessing Redis, as simple as that.
However, because of this popularity, the applications started running into the scalability constraint, where the single Redis node was not able to serve the workloads anymore. Either sometimes there's too many client connections, too many operations per second hitting the vertical scaling limits of a single node. There was no native solution available at the time. Then this led to new architectures, which I refer to as the proxied architectures or the gateway architectures, where you basically take the independent Redis nodes and then you stitch them together. You make it look like one giant system, one giant cluster. Your data is sharded across them, and your application basically connects with them. The blueprint, what I have, is the legacy applications and then a gateway layer, and then the Redis servers backing it.
What we see in 2009, over time, so it was 2015, 6 years later when Redis Cluster was released, Redis Cluster improved upon these sharding capabilities. It provided a native sharding capability built into Redis, where a bunch of nodes could work together to provide that one giant cluster view to the end client. Then, the proxies were still there, legacy apps were still there but the trend continued. For one, the Redis Cluster offering was not mature enough. It was new. It was complex. You needed some complicated logic on the client, so not everyone got it right. It was not gaining that much popularity, which doubled down on the proxied architectures. They continued to grow. The legacy applications, even if they wanted to adopt a clustered offering, they couldn't because they didn't want to change the client, so what they did is they, again, continued with the proxy layer in the middle.
You put the cluster behind the proxies, and then you can make it work. You don't have to do the independent nodes behind the scenes. You can still use clusters, and then the proxies would abstract it away from the end client application. Legacy constraints was feeding into this growth, or is still feeding into the growth. The second one was related to simplicity, or abstracting out, or decoupling your client applications from the server, so that's like another one of the non-functional reasons why you see proxies coming into picture. The third one is more technical, I would say, less of a constraint from a legacy perspective. This has to do with connection multiplexing where your application scale is so high, and there are very few applications that I've seen at that kind of scale where you are basically running with 100,000 pods, the client pods. They are trying to connect to the cluster behind the scenes.
If you were to connect to the Redis, Valkey cluster, whatever, directly with those 100,000 connections, at least one each coming from the client, that would definitely kill the cluster. In those cases, you want to have an architectural improvement using proxies, for example, which can take on or offload the connection management. What happens is you get 100,000 connections onto the proxies. They are stateless, so you can scale them out horizontally. Then they would take those requests and multiplex them onto fewer connections against the backend server. The idea is to take away the connection handling responsibilities, move them to the proxies. These are some of the legit or the real requirements around where proxies have served a purpose or continues to serve the purpose.
In this architecture, we use a common open-source Envoy proxy. There are 5 to 10 proxies which are widely used in the industry. I just took one of them, but it has nothing specific to do with the Envoy proxy, so we'll use this example throughout our discussion here. In this example, you have the Envoy proxy layer, which is hosting this logic. The logic could be in two parts. One is, you have independent Redis, Valkey nodes behind the scenes. I've switched to Valkey here instead of Redis. The Envoy could shard the client data using a simple hashing scheme. It's not a database, so there are a little bit of data losing concerns, but it can make it work. Then your client connects to Envoy proxy. It can talk to any Envoy proxy. Then it routes the request to the right node. The other mode is the cluster mode, where Valkey is self-aware of how the data is sharded, and Envoy proxy can also connect to that. That's the advantage of these proxies, that they can speak multiple protocols with the underlying backing system.
We talked about the journey of Redis, how proxies came into picture, how they continue to exist, some of the top reasons behind these proxied architectures. Now we're going to dive into the real-world tradeoffs or the metrics around price, performance, latency for these proxied architectures. What do they actually look like? Let's start with the latency. One could categorize it as a performance dimension. Latency is one way to look at it. Coming to our architecture with three Envoy nodes, three Valkey nodes. Now, to study the latency aspects, the performance aspects, I'm going to simplify this a bit. We are going to just reduce the three Envoys to one, and similarly for Valkey, three to one. It's a pretty simple setup with the client talking to the proxies, proxies talking to the Valkey node behind the scenes. The run I have here, it uses a utility called memtier_benchmark.
Nothing special here. It's just a tool to drive load, and basically measure some characteristics of the system. I have this deployed on EC2, AWS. You can see the VM types. Both of the proxy and Valkey are the exact same machine types. They are in the same availability zones, no cross-zone hop, wanted to reduce the interference. They are both 8-core machines each. This is a benchmark run from my terminal. Don't worry about the grayed-out details. What's important is we talked about the memtier_benchmark. Then the second aspect is it's a read-heavy workload, which you see typically in practice. In this case, it's roughly a 90-10 ratio. More than 90% of the traffic is read traffic. In this setup, what you see is this setup is driving half a million QPS. Half a million queries per second is being driven, which is roughly 90-plus percent reads, and then the remaining is sets.
The latency. Since we wanted to talk about the latency. For these half a million queries per second, the latency that you get, so we measure on two sets, one is the tail, the p99. That's roughly 2.5 milliseconds. This is in the milliseconds category where it could go from anywhere like 1 to 10 milliseconds, roughly speaking. Then there is p50, which is your median latency, and this is close to 1 millisecond. This is what I would refer to as our milliseconds app. These numbers are still super good. Like, they are driving half a million QPS. The latency is super low, single-digit millisecond. Not bad at all.
For each of these setups, we are going to look at a little bit deeper into the latency, price, performance aspects one by one. First, we are going to look deeper into this p50 latency. I like sometimes appreciating the system characteristics, like how it ends up to be where it is right now. In case of the p50 latency, we talked about 1 millisecond, and as you notice, this latency is measured from the client side, so it's a full round trip covered in 1 millisecond. There are two hops going on. These two are network hops, so in order to measure the cost of a network hop, I used a utility, My Traceroute. It's a Linux utility. It's similar to a ping, just has more fancy knobs around it. What you see here is between the client and the proxy node, there is a 300 microseconds hop. This 300 microseconds hop is basically in the same availability zone.
Like I said, not much of an interference, not a cross-zone hop, which gets to be higher. The same 300 microseconds would go from Envoy to Valkey, all round trip. Ultimately, what it ends up to be is 600 microseconds, the two hops, 300, 300, out of the p50 of 1 millisecond gets spent on the network. What you're left with is 400 microseconds, and that 400 microseconds involves the full proxy processing to route the request. It involves the Valkey server processing to serve the dataset. Then it also involves a little bit of client processing, where the client sends the request, gets the response, measures the time taken. All that stuff is basically done in just 400 microseconds. Network, maybe you can't do much about it for a remote lookup, but that's how the latency breakdown looks like for a p50 latency.
Next, we will look at how much do you need to spend to get a certain amount of throughput, measured in queries per second. Cost is important. As engineers, we generally talk about latency, performance. I like to also talk about the dollar. We're going to dive into that. In our setup, recall that we got half a million QPS at 2.5 millisecond tail latency, pretty good. The CPU measurement during this run was 90% CPU utilization on Envoy, which is pretty high, so it means it was almost completely taken. Then the Valkey server CPU utilization is basically 60%. There still seems to be a lot more capacity which is remaining on Valkey to be utilized. Just mapping it back to the diagram, visually you can see there is 90%, 60%. Why is this proxy 90%? Why is it so compute intensive? Why is it consuming so much CPU?
In the caching world, this is where you'll see that it's quite different from the databases world, where we worry usually less about the QPS. It's more about latency, the I/O, and stuff like that. In the caching world, there are two main things that would typically come up. One is the QPS. Sometimes the applications are offloading the QPS from the database to the cache. Generally, caching applications, they would have high QPS, or they have a lot of in-memory storage. You are paying for in-memory storage, which is much more expensive than our typical disk-backed storage. This is where the caching workloads are interesting, because you have to pay attention to the QPS aspects as well. Now, because of this QPS, this traffic is being forwarded through the proxies. What I want to do is, I want to build an intuition around where is this compute going. The 90% consumption on the Envoy, what the breakdown looks like.
How we talked about the latency, where we talked about the network hops, network hops taking majority of the latency. In case of the CPU, so one intuition is basically, when you're forwarding traffic through the proxies, it's doing twice the amount of I/O as compared to the underlying Valkey server. If you look at the Valkey server, it gets a request, sends a response out. In case of the proxy, it's a pretty simple thing. Client sends a request, that's one. Request goes out to the Valkey server. The response comes back from the Valkey server, that's the third one. Then the response goes back to the client. The proxy node has to move this data or the underlying network packets twice the time of what a typical I/O would look like for Valkey. This whole process of moving the network packets is very compute intensive. What you see here is the process routing the request.
Underneath, the network card, once it receives the packet, it has to put that into the host memory because we are running virtual machines, so they are running on a physical host. Once it gets into the memory of the physical machine, another process would route it to the guest VM, the guest OS. Then from there, through interrupts or some mechanism, it would get delivered to the kernel's TCP stack, which is where eventually it would find its way to your application where your code of routing the request runs. Then it would be the reverse path when it has to send something out.
This was about building an intuition. They're twice the I/O. Now let's look at the real data. We want to validate our intuition now with a profiler. Flame graphs are also my favorite, but they're also very hard to read when it comes to slides. Flame graph is just a visualization tool which captures, during the actual run, the stack frames, and then sees what stack frames appear more often than the others. Then if something is appearing more often than the other, you can take enough samples to determine maybe this is the area where it spends more CPU than the others. Don't worry about reading these stack frames. I'm just going to do a higher-level analysis here. We looked at the 90% Envoy CPU consumption. I have a breakdown, which you can also measure from the process. There is user mode, there is system mode. User mode is where your application is running.
The code written, whoever built the proxy would run into the user mode. Then it interacts with the system calls, the TCP stack, that's all basically kernel mode. The split of this 90% is very even, like 45%, 45%, user mode versus kernel mode. Lot of I/O, there is obviously a lot going on with the application. You see the two big boxes I have, these are parts of the flame graph. One part on the left is about the application processing. This is less about the I/O, but it still needs a lot of CPU. This can be heavily optimized, writing different code and all that. The right part is the TCP send and receive path which involves a little bit of application and also the kernel frames. It's like the I/O, what we said is twice as expensive. The I/O does form a major chunk of this whole 90% CPU utilization. The key point we are trying to understand is when you say, proxy consumes a lot of CPU, what is it actually doing? It's actually routing requests. A significant amount of CPU spent there. The rest of the CPU, in this case for Envoy proxy, it's half, it's spent in the I/O.
What we have established is this 90% Envoy, given that the Valkey server has capacity. To me it reads like, Envoy is like the chokepoint here, and then maybe there is something we can do. The idea is to, now I'm going to add one more Envoy proxy to the setup to horizontally scale this layer. Then, using these two Envoy proxies, they are stateless, so I'm going to load balance the request. The client would just pick any one arbitrarily and drive the traffic. Then we're going to just repeat the benchmark. With these two proxies fronting the server, we have 1 million QPS driven from this system. When we were using one Envoy proxy, we had half a million QPS being funneled through the system. Now with two, we have 1 million. That's good because it has a nice scaling characteristic. One was half, it doubled when we doubled the hardware there.
That's good. The latencies are quite similar. p50, again, 1 millisecond. Tail latency, 2+ milliseconds, similar to what we had before. We still retain our milliseconds app roughly. Then we were able to hit 1 million QPS out of this system setup. Since we talked about the dollar cost, let's look at what this setup would cost us. The eighth generation Graviton instance that I'm using costs about $230 a month. The hardware we have provisioned is about $400 roughly for proxies, $200-something for Valkey. This comes out to be about $700 a month to deliver on 1 million QPS at low single-digit millisecond tail latency. This is an important result. Let's take a moment to capture this. I'm going to come back to this later in the discussion. This is not bad. If I gloss over the architectural details, you have an in-memory setup. In-memory is expensive. Everyone knows that.
We are spending $700, justified for the business. Nothing seems wrong. Note that the $700 in reality is actually a lot more. It's going to be close to easily $1,500 or something, because this does not have high availability built in. With high availability, you will have copies and then you'll have cross-zone hops. There will be a lot of network jitter, more cost involved. What that would lead is like worse latencies, higher cost. What we are talking is a little bit of a crafted setup in the same availability zone without those variables.
What we covered so far, we talked about the latency aspects. We talked about the cost associated with delivering 1 million QPS at a certain latency. Typically, architectural evaluations would stop here because we added the proxy. We were analyzing the proxied architecture. It's a stateless system. Stateless systems are usually harmless, from our experiences. Very few times we would look at anything beyond the setup. The third piece is reliability. It has an interesting pitfall associated with it, which is a single point of failure. Our setup has a single point of failure. At least to me, it was not obvious when I first exposed to this problem as to, what is special about this architecture where reliability could take a hit. It's related to head-of-line blocking. It's related to single point of failure. We'll go over that. To understand these pitfalls, we want to drive some traffic. Again, nothing fancy.
We have our architecture. I'm going to let go the single node one and bring back the three proxies, three Valkey servers because more nodes are interesting. We're going to drive some traffic. I have a custom script, not the benchmarking one this time, just to capture the client-side availability, what it looks like. It's driving a small amount of traffic, and the availability, what you see here is 100% every second, which means all the 100 QPS, basically all the requests are successful. Nothing comes without failures, or even worse, slowness. To make things interesting, we're going to take a shard and inject some slowness into it and see how the system behaves. This is Lua scripting. Redis has this powerful feature, Valkey has it as well. It basically is like a server-side function where you can inject some logic and make it run on the server. The code here is a workaround to basically make the main server slow down for about 5 seconds, so just to induce a slowness or a failure there. You can reproduce this easily. Just don't do it in production. It's pretty deadly to the production server.
Now let's see what happens to our client-side availability. We have the 100% run rate, and then for those 4 to 5 seconds when we basically injected the slowness, our availability dropped to zero. The client is experiencing full unavailability, no requests going through for those 5 seconds when we inject this failure. We had evenly distributed traffic in our setup. What even distribution means, like our shard_1, which was experiencing this slowness, was getting one-third of the traffic. The other shards, we did not touch them at all in production or in our setup. They were getting two-thirds of the traffic, but yet our net availability is zero. As you can also see, all my traffic is now routed to shard_1. Shard_2 and shard_3 are actually not even getting any traffic. That hints us to something, like they are all healthy, no CPU consumed, like what is going on?
Imagine if you were running a real server, instead of the three nodes, you had tens or even hundreds of nodes, both for the proxy setup and the shard, one shard going down, taking down your entire cluster. We're going to root cause this. Since the client reported zero availability, so I'm going to just zoom in to the client to understand what is going on. This is what the healthy state looks like. We follow the best practices. The client has to connect to the proxies. You need connections, so you need a connection pool. Do not create connections all the time, take that overhead. This connection pool is capped to five connections. Then these five connections are basically serving the request for the shard_1 or whatever request comes up. Anything more would get throttled. Requests to the shard_1 start to take longer. Because of this, the request to the other shards, they get starved.
Now the entire connection pool becomes occupied with shard_1 requests, because they were slow, so they were not freeing the connection sooner. This is what leads to the outage. As you can see, all my requests to shard_1 have taken over the connection pool on the client side. Shard_2 and shard_3 are not getting any requests. They're all being rejected. This is how a single shard can turn the entire system into a full outage. To summarize what we learned with these proxied architectures. We spent $700 to achieve 1 million QPS with low tail latency. Then we also saw how one shard crashed the whole system.
Part 3: The Microsecond Playbook - The Strategic Path to µs Efficiency and Resilience
Our last part is about microseconds. Let's get deeper into the microseconds architecture. There is a quote from Dr. Werner Vogels, which reminds me, we always say that we have done it this way, so we would want to look for ways to question the current system, to constantly optimize and improve our systems. Just like the space shuttle where we got rid of the wings by questioning the runway landing requirement, getting to the capsules which were much more efficient, I'm going to do the same exercise with these proxies. I'm going to remove these proxies. Maybe they are critical to some use cases where you would definitely need them. For our discussion, I'm going to remove them and see what happens. After the proxies are gone, our clients are able to access the servers directly, but you have multiple servers, so you need to know where the data is.
What you need is a smart client which can understand the topology. The underlying Valkey cluster has a way for the client to learn the topology and route the requests accordingly. What we are going to do is we are going to do an exercise of the latency, price, performance, reliability aspects, but in the reverse order. We're going to start with reliability, then talk about cost, and then finally talk about latency. How we did with the proxied architecture, in the same direct access architecture, we are going to inject slowness in this shard. Use the same script, exactly the same setup. This time what you see is that availability is not zero, like earlier. It's partial outage. It hovers between 60% to 80%, which roughly is two-third of our traffic as expected. When shard_1 failed, my availability reported is 66%. Now let's zoom in into the client, like how we did with the root causing to understand what is going on here.
When your client is connecting directly to the shards, it's as simple as you have different connection pools, one per endpoint. Each connection pool has a bunch of connections to send requests to that shard. Then what happens is when shard_1 gets to slow down, it occupies all the connections in that pool, in the respective pool, and then more requests come up, even more requests, they start to get throttled. What's important is the shard_2 and shard_3's requests are all protected, they are isolated. The fault isolation, what you have on the shard, because the shards are independent of each other, now extends to the client side where the other requests are not impacted. This reminds me of a bulkhead pattern. This is a pattern, goes back to ancient Greece times with ships where they built these bulkheads. They're still there in practice in different shapes and forms. These bulkheads are basically like compartments. They are for the structural integrity of the ship. One other important aspect they serve is fault isolation, similar to our sharded architecture where one part gets damaged, it's not like it will sink the entire ship. It has some impact, but the ship still continues to float.
We just covered the reliability where we saw how with the proxied architecture, one slow shard takes down the whole cluster, and how in the direct access architecture we were able to isolate that fault as expected. Next on our list is price, performance. For price, performance, we have a single Valkey shard, no proxies. We are going to run our benchmark, which is able to achieve 1 million QPS. This Valkey server is able to achieve 1 million QPS without any proxies. The Valkey server just costs $230, like what we are basically able to drive here. This is our latency numbers with our Valkey setup. We have about 567 microseconds at the tail latency instead of the 2+ milliseconds. This is at 1 million QPS.
Summary and References
To summarize and put it all together, we have the cost, which dropped to one-third without the proxies to deliver the same 1 million QPS. We have four times higher latencies with the proxies set up at the tail. Instead of 600 microseconds with Valkey server, we have about 2.4 milliseconds. Here are the references I used for this talk. We basically covered the space shuttle, what you'll see here, Dr. Werner Vogel's Laws of Frugal Architect, and the Valkey's 1 billion RPS innovation. I didn't get to go deeper there, but you can read about it with these references.
Key Learnings
The last slide is about the key takeaways, how we went into the direct access architecture where we understood how the network hop, the 300 microseconds was taking out almost all the latency out of the 360 microseconds for p50. We saw how the architecture impacts the cost, the $700 versus just the $230. How a single point of failure in the proxied architecture goes away with the direct access architecture. Finally, Valkey server is so efficient because of all the community contributions behind the scenes which help us drive this infrastructure cost to be lower.
See more presentations with transcripts
Recorded at:
Aug 06, 2026
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み