リアルタイム市場向け大規模 AI セーフティシステム構築の発表:SafeChat の実践
本文の状態
日本語全文を表示中
詳細モードで約37分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
InfoQ AI/ML
コストの高い LLM のみを用いる従来のパイプラインを、明らかなケースをフィルタリングする高速内部モデルと、微妙な判断を LLM が多軸スコアリングで行う構成へ置き換えた。
AI深層分析を開く2026年8月22日 20:37
AI深層分析
キーポイント
ハイブリッドアーキテクチャの採用
コストの高い LLM のみを用いる従来のパイプラインを、明らかなケースをフィルタリングする高速内部モデルと、微妙な判断を LLM が多軸スコアリングで行う構成へ置き換えた。
コンテンツ非依存プラットフォームの構築
特定のカテゴリに縛られない汎用的な AI モデレーションプラットフォームを DoorDash が開発し、リアルタイムマーケットプレイスでの運用を開始した。
バックテストとノーコードワークフロー
システムの改善を検証するためのバックテスト機能と、技術的なコーディングを必要としないノーコードワークフローを導入して開発効率を高めた。
既存システムの破棄と再構築
実環境で機能していた安全システムを捨てて、より強力な新システム「SafeChat」を構築した。
ブラジル拠点のチーム体制
信頼と安全性(trust and safety)エンジニアリングチームをサンパウロ拠点から率いている。
重要な引用
replacing costly LLM-only pipelines with a hybrid pattern: using fast internal models to filter obvious cases, LLM multi-axis scoring for nuanced decisions
no-code workflows with backtesting
cut safety incidents while scaling to millions of daily messages
By the end, I want to leave you with two things, how we built SafeChat at DoorDash, and one architectural pattern that you can take home and use at almost any AI use case that you have.
編集コメントを表示
編集コメント
LLM の運用コストが課題となる中、フィルタリングと判断を分担するハイブリッド手法は実務的な解決策として極めて示唆に富んでいる。QCon AI で発表されたこの知見は、大規模システムにおける安全対策の設計指針として広く参照されるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
InfoQ ホームページ
プレゼンテーション一覧
SafeChat: リアルタイムマーケットプレイスで大規模に構築する AI 駆動の安全システム
プレゼンテーションを見る
再生時間:
ダウンロード
42:22
画像/presentations/doordash-llm-ai-moderation-platform/en/slides/slide-1786348150151.jpg)
概要
DoorDash のソフトウェアエンジニア、Bruna Pereira は、コンテンツに依存しない AI モデレーションプラットフォームの構築について解説します。高コストな LLM(大規模言語モデル)単独のパイプラインから、ハイブリッド型のアプローチへと移行しました。具体的には、明らかなケースを高速でフィルタリングするために内部モデルを活用し、微妙な判断が必要なケースでは LLM による多角的スコアリングを採用しています。また、ノーコードワークフローとバックテスト機能を組み合わせることで、安全インシデントを削減しながら、1 日数百万件のメッセージに対応できるスケーラビリティを実現しました。
プロフィール
Bruna Pereira は DoorDash のソフトウェアエンジニアです。10 年以上にわたりソフトウェアエンジニアリングに従事し、難題の解決や大規模システムへの取り組み、本番環境での構築を通じた学習を得意としています。
コンファレンスについて
QCon AI は、これらのワークロードを安全にスケールさせるために必要なエンジニアリング分野に特化した、実践者主導のイベントです。同業他社が生産環境で採用しているアーキテクチャのプレイブックや、失敗事例のメトリクスへの直接アクセスを提供します。
INFOQ EVENTS
2026 年 8 月 27 日午後 1 時(EDT)## スケールする AI エージェントが失敗する理由:コンテキストはデータレイヤーの問題だ 登壇者:Boyd Stowe氏(Tacnode 創設ソリューションアーキテクト)
2026 年 9 月 17 日午後 1 時(EDT)## PR の向こう側:エージェント型ソフトウェアデリバリーの新たなコントロールプレーン 登壇者:Mohit Suman氏(Harness スタッフプロダクトマネージャー)
トランスクリプト
Bruna Pereira: こんにちは、ブラーナ・ペレイラです。DoorDash のソフトウェアエンジニアとして、本日は実際に運用現場で直面した安全課題を解決するために構築したシステムについてお話しします。その仕組みが機能し始めた瞬間に、私たちはあえてすべてを捨てて、さらに強力なものをゼロから作り直しました。
このセッションの最後には、2 つのポイントをお伝えしたいと考えています。まずは DoorDash で「SafeChat」をどのように構築したか、そして次に、あらゆる AI ユースケースで即座に活用できるアーキテクチャパターンです。
私は現在、DoorDash でソフトウェアエンジニアとして 1 年半勤務しており、ブラジルのサンパウロ拠点にある信頼と安全(Trust and Safety)のエンジニアリングチームを率いています。その前には 3 年間、自社スタートアップで事業を展開していました。私の専門分野は主に FinTech 領域で、大規模システム構築の経験が豊富です。
DoorDash マーケットプレイス(安全性最優先)
DoorDash はマーケットプレイスです。消費者が注文できるアプリ、食材を届けるダッシャー(配送員)、料理を用意する店舗という 3 つのグループが存在し、彼らは移動の異なる段階で出会います。チャットを通じてやり取りをしたり、配送中の電話通話を行ったり、時には直接対面することもあります。
DoorDash にとって重要なのは、人々が実際に安全であることと、プラットフォーム上で安心感を感じてもらうことの両方です。この 2 つは、当社にとって同等の製品指標として扱われています。
安全について語る際、プラットフォーム上で発生するインシデントの多くが、言葉による虐待に関連しています。これらのグループはチャットや音声通話を通じて接するため、手段の違いは問題になりません。不審な兆候を感じた瞬間に即座に対応する必要があります。ソーシャルメディアプラットフォームとは異なり、当事者間に強い信頼関係を築く機会がないためです。
これらのやり取りは、最長で 60 分、最短でも 40 分程度続きます。何か危険な兆候が見られたら、即座に対応する必要があります。
スケーラビリティの観点から言うと、チャットでは毎日 400 万件以上のメッセージが交換されています。配送中にも消費者と Dashers(配達員)の間で 40 万件を超える通話が行われ、チャットや SMS を通じて 20 万件以上の画像がやり取りされています。
特にチャットにおいては、「すべてのメッセージが安全であると判定されること」をゴールにしています。これがこのシステムの難しさを際立たせる要因です。膨大な数のメッセージを処理する必要があるからです。チャットの文脈では、メッセージが安全かどうかを判断するのは数分の一秒の世界です。そうしなければ、ユーザーのチャット体験を損なうことになります。
当初、この課題について議論した際、特にビジネス側の最初の反応は「LLM(大規模言語モデル)を使えば解決する」というものでした。「LLM を呼び出して、メッセージが安全かどうかを判定すればいい」。安全なら配信し、不安全ならブロックする。理論的にはこれで機能するように思えます。
しかし、実際の運用では、私たちが扱っているような膨大な量に対しては通用しません。もしあなたが以前に本番環境で LLM を使った経験があればご存知の通り、LLM 呼び出しのレイテンシ(応答遅延)は大きく変動します。私たちのユースケースでは、1 回の呼び出しあたりの平均レイテンシが 2〜10 秒にも及びます。
さらに、1 日に 400 万回も LLM を呼び出す場合の費用を想像してみてください。これは現実的な解決策ではありません。
では、LLM を賢く活用するにはどうすればよいのか。まず着手したのは、自社のデータを理解することでした。
数ヶ月をかけて、「自社における不安全さとは何か」を定義し、実際にメッセージの何%が不安全なのかを把握しました。多くのメッセージが不安全だと想定していましたが、その割合はどれほどでしょうか。ここで即座に決断を下すことが目的ではありません。保有するデータから学び取ることに注力したのです。
チャットログに計測機能を組み込み、市場で利用可能な無料のモデレーション API を活用しました。非同期でこの API を呼び出し、不安全なメッセージのカテゴリを特定します。誰もが飛ばしたがる地味なステップですが、これが重要です。社内では「行動に移す前に数ヶ月かけてデータを理解すべきだ」と主張しています。モデルが安価か高価かに関わらず、データを知っていなければ何も得られないからです。
この分析を実行した結果、既知の事実を確認できました。メッセージの大部分は安全ですが、具体的な数字を算出できたのです。不安全なメッセージはたったの一桁%に過ぎませんでした。この発見がアーキテクチャ全体を決定づけたのです。「ほぼ常に正しく動作する、コストを抑えた積極的なレイヤー」を構築できることが明確になったからです。
その後、小型の分類器を構築しました。これは前工程で収集したデータを用いて訓練された機械学習モデルです。
このモデルには3つの役割がありました。まず、90%の確率で100ミリ秒以内に応答する高速性。次に、呼び出しごとのコストが発生せず、インフラ運用コストのみで済む低コスト性。そして、「明らかに安全な内容」を正確に識別できる特化性です。
この層は最終的な判断を下すものではありません。米国の空港の金属探知機を想像してください。金属探知機は「あなたが脅威である」と宣告するのではなく、「より詳しく調べる必要がある対象だ」と示すだけです。
この自社開発モデルで安全と判定できないメッセージについては、LLM(大規模言語モデル)に処理を委ねます。ただし、対象となるのは全メッセージの10%未満です。ここが最も重要な設計上の選択となります。
LLMを呼び出す際、「このメッセージは安全か」という単純な質問はしません。代わりに、複数の軸に沿って分類やスコアリングを行います。後ほど数スライドでその重要性をお話ししますが、文脈に応じて「どの程度脅威的か」「どの程度卑猥か」「どの程度性的か」などを問います。
この層こそがモデルの知性を発揮する場所です。なぜなら、ここでは難しい質問にのみ答えさせるからです。
全体的なアーキテクチャについてお話ししましょう。メッセージが着信すると、まずノイズを除去します。空のメッセージや画像添付ファイルは除外し、これらは別のパイプラインへ送られます。また、一般的な挨拶のような定型表現も取り除きます。
その後、最初のレイヤーである小規模なモデル(クラスフィア)に送られます。メッセージが安全と判定されれば、そのまま配信されます。問題がある場合は、第 2 レイヤーの LLM(大規模言語モデル)へ回されます。LLM は複数の軸に沿って分類を行い、その結果に基づいて段階的な対応措置を講じます。
これが本講演全体で示す仕組みです。多くのメッセージが通る安価なレイヤー、コストのかかるレイヤー、そして最終的なアクションという流れです。各レイヤーは得意とする業務を適切なコストで担っています。この「適材適所」の考え方が最も重要なポイントであり、もしあなたがモデレーションシステムを構築する際には、ぜひ覚えておいてほしいスライドです。
LLM に「安全か不安全か」という Boolean(真偽値)を問うのではなく、スコアを求めましょう。その理由は、Boolean が単なるフラグであるのに対し、スコアは調整可能なノブのような役割を果たすからです。スコアを採用すれば、得られた数値に応じて段階的な対応が可能になります。また、後から新しいカテゴリを追加する際にも、すべてを作り直す必要はありません。閾値が高すぎる、あるいは低すぎると感じれば、柔軟に移動させることもできます。
さらに、このアプローチは LLM が得意とする領域に合致しています。例えば、「このメッセージは安全か不安全か」と問われた場合、文脈によって「何が不安全なのか」の定義が異なり、回答もバラついてしまう可能性があります。その代わりに、「安全性とは何か」を理解した上で、LLM にその軸に沿ってスコアリングを依頼しましょう。そうすれば、結果ははるかに安定したものになります。
このスコアリングが重要なのは、メッセージの深刻度に応じて適切な対応を行う必要があるからです。当社の文脈では、低レベルの違反は悪口や汚言などです。このような場合はメッセージをフィルタリングしつつ、送信自体は許可します。
中程度の違反(例えば相手を侮辱した場合)では、メッセージを完全にブロックして配信させません。
高レベルの違反(脅迫など)の場合は、メッセージをブロックするとともに、被害を受けたユーザーに対して注文のキャンセルと全額返金を提案できます。
極めて深刻な違反の場合には、注文キャンセル、メッセージブロック、加害者への警告、そして被害者をシステムから完全に排除するという対応を行います。これらの多様な対応が可能なのは、まさにこのスコアリングがあるからです。もしスコアがなければ、「失礼な表現」と「危険な行為」を区別することができなくなります。
ここではチャットについて説明しましたが、音声や画像の場合も基本的な構造は同じです。
スコアリングエンジン自体は共通です。異なる点は、画像の場合は安価な内部レイヤーの代わりに、暴力やわいせつなコンテンツ、危険な画像をすでに十分に見分けられる商用ビジョン API を採用し、これを最初のフィルタとして活用しています。
音声については少し複雑です。チャットであればメッセージを読み込んでから配信するか否かを判断できますが、音声の場合、メッセージを文字起こしして分析する時点で、相手はすでにその言葉を聞いてしまっています。一度届いてしまったメッセージを完全に阻止する方法はありません。私たちができるのは、対応の仕方を変えることです。不審な兆候を検知した瞬間に通話を切断したり、もちろん、注文のキャンセルを提案したり、自社の側から注文を取り消したりすることが可能です。
結果
この対策を実施した後、SafeChat と呼ぶ私たちの仕組みを導入する前後で、音声による虐待に関連するインシデントがどれだけ減少したかを測定しました。その結果、約50%の削減を実現できました。これはモデルの精度向上のような数値ではなく、人間への実際の被害を減らしたという実証です。この数字こそが、データを理解し、顧客を支える強力なシステムを構築するために費やした数ヶ月の努力正当化する根拠となります。
コンテンツ非依存のモデレーションプラットフォーム - 構成要素
この発表を終え、製品をリリースし、チームが評価され、全員が帰宅する——そんな理想の結末を想像していました。しかし実際には、私たちはすべてを捨ててしまいました。ただし、得られた知見や学習成果、訓練したモデル、収集したデータまでを廃棄したわけではありません。私たちが手放したのは、システムそのもの、つまり「SafeChat」という仕組みです。なぜそう決断したのか、その理由をお話しします。
SafeChat の展開が始まった当初、DoorDash 内の他チームや外部のチームから、「素晴らしい取り組みだ。自分たちのユースケースでも同様に実施したい」といった声が寄せられました。具体的には、「配達員や利用者のプロフィール画像を審査できないか」「登録時の名前をチェックできないか」「食品レビューのモデレーションは可能か」など、さまざまな要望が相次ぎました。さらに、安全性とは直接関係ない領域にも関心が向けられ、「チャット内や電話通話での不正行為を検知できないか」といった提案も出されました。
これらの要件を一つずつ実装し始めると、毎回ゼロから作り直さなければならなくなることに気づきました。しかし、彼らが本当に必要としていたのは SafeChat 自体でも、私たちが構築したシステムそのものでもありませんでした。彼らが欲しかったのは、私たちが生み出した「パターン」です。
具体的には、「最初は安価に、次に高価に、そして段階的にアクションを実行する」という仕組みです。これが彼らの真のニーズでした。そこで私たちは、これをあらゆるユースケースで誰でも利用できるプラットフォームへと変換できないかと考えました。その結果、コンテンツ非依存のモデレーションプラットフォームを構築しました。
このアプローチの核心は、コンテンツの意味やビジネスロジックを理解する必要がない点にあります。私たちが把握し、実行すべきことは、意思決定のログを記録する方法、異なるモデルプロバイダーと統合する仕組み、そして条件に応じてステップを切り替える方法だけです。意味や文脈は各チームが持ち込み、プラットフォームはその調整役(オーケストレーション)を担当します。
私たちはチームに対して、「コードを書かずにこれを実現できる」と約束しました。彼らは UI を介してプラットフォームにアクセスし、設定を行うだけで、必要なアクションを実行できます。これがこのプラットフォームの最大の強みです。
コンテンツガイドラインがある場合でも、不正行為対策や安全性向上のためのユースケースであっても、関係ありません。どのケースでも対応可能です。その実現を可能にする基盤となるコンポーネントがプラットフォームには用意されており、これらについてご紹介します。
現在、このプラットフォームには3種類のモデルが用意されています。
まず「内部モデル」です。これは自社サーバーでトレーニング、ファインチューニング、デプロイが可能で、完全に自前で管理できるモデルです。
次に「外部モデル」。ベンダーとの契約を通じて利用するモデルで、外部から提供されるプロンプトもこれに含まれます。外部プロンプトとは、市場にあるあらゆるベンダーのLLM(大規模言語モデル)で使用できる汎用的なプロンプトのことです。
内部モデルは自社内でトレーニングされ、ホストされます。MLプラットフォームがこれを支えており、市場で利用可能なあらゆるモデルを活用しつつ、自社のラベル付きデータを用いた学習を可能にします。ラベル付きデータがあれば、それを活用してモデルを訓練し、自社のインフラストラクチャ上でデプロイできます。
これらの内部モデルは、プラットフォーム内で利用可能なAPIを通じて提供されます。その目的は、低コストかつ高速な運用を実現することです。
学習時には、事前に定義した入力・出力スキーマを設け、このスキーマに従うことですべてのクライアントが利用可能になります。SafeChatのために構築した小型の分類器は現在では内部モデルとして運用されていますが、外部モデルも併用しています。
ここで重要なのは、市場に既に存在する機能を無理に再開発しないという考え方です。例えば画像安全性の確保はすでに解決済みの課題であり、ゼロから作り直す必要はありません。プラットフォームの理念は、単なる理由で既存の機能を再構築することではありません。必要な機能は外部に用意されており、提携契約を結んでおり、プラットフォーム上で一度統合すれば、どのクライアントも自由に利用できます。
課金の追跡は異なる API キーによって行います。また、外部プロンプトも活用しています。これらは最も柔軟なアプローチであり、LLM ゲートウェイを通じて市場にあるあらゆるベンダーのモデルと統合可能です。
もし、この LLM ゲートウェイの構築プロセスについてさらに詳しく知りたい場合は、DoorDash のチームによる「DoorDash での LLM ゲートウェイ構築」に関する登壇資料が参考になります。ここで紹介されているゲートウェイは、プラットフォームと外部のあらゆるプロバイダーから利用可能なモデルとの間に位置し、仲介役を果たします。
クライアント側では、対象とするモデルを自由に選択できます。また、各プロンプトに対して入力スキーマと出力スキーマも定義可能です。この LLM ゲートウェイが提供する機能を活用すれば、コードを書き直すことなく、フォールバックやリトライ戦略を宣言的に設定できます。
具体的には、LLM モデルを利用する際によくある問題——応答がない、プロバイダー全体がダウンしている、あるいはモデルが期待通りに動作しない——に対処するために、この仕組みを使います。もし対象としたモデルが利用不可になった場合でも、別のプロンプトをフォールバックとして設定しておけば、システムが復旧するまで代替手段を提供できます。
さらに、LLM からの呼び出しに対して望む出力スキーマ(JSON 形式など)も定義可能です。ただし、LLM がその形式で応答しないケースも想定されます。そのような場合でも、クライアントに構造化されたエラーを返すまでのリトライ回数を事前に設定しておくことで、堅牢な対応が可能になります。
Composing Moderation Agents
これらが基本構成要素です。これらのブロックを組み合わせて、私たちは「モデレーションエージェント」を構築できます。モデレーションエージェントとは、ステップを柔軟に組み合わせ可能なパイプラインまたはワークフローのことです。
例えば、内部モデルの出力に基づいて条件分岐を行うケースがあります。この条件は、内部モデルが返した結果に応じて評価されます。条件を満たせば特定のステップへ進み、満たさなければ別のステップへ進むといった制御が可能です。ここでは、外部ベンダーへの転送や LLM プロンプトの実行などが選べます。
ただし、フローは自由にカスタマイズ可能です。最終的なアクションも実行できますが、その責任と決定権はクライアントにあります。各ステップの出力結果を把握した上で、必要なアクションを実行するか、あるいは何もしないかを選択できます。
条件式の定義については、これもすべて UI 上から直感的に行えるよう設計されています。
SafeChat のパイプラインでは、前段のモデルからの出力を用いて、シンプルなものから複雑なものまであらゆる条件を記述・表現できます。この仕組みは非常にシンプルです。
内部モデルが返す不審ラベルのスコアが 0.5 を超える場合、その結果は LLM プロンプトへ送られます。LLM は各カテゴリとそのスコアのセットとして応答し、その後アクションが実行されます。一方、スコアが 0.5 以下の場合は即座にアクションへ進みます。
ここでは特定のアクションを実行しないケースもありますが、実際にはメッセージ内に悪口が含まれている場合は検閲(フィルタリング)が行われます。どのステップ間でも任意の条件を定義できるため、独自のモデレーションエージェントを構築可能です。このエージェントは主に 2 つの種類に分けられます。1 つ目は同期型モデレーションエージェントで、これは HTTP リクエストとして実装されます。
モデレーションエージェントの実行中は接続を維持します。エージェント内で定義されたすべてのステップを実行しますが、意思決定のゲートとして必要な場合のみ同期処理を使用します。例えばチャットでは、メッセージが安全でない場合にブロックできるようにする必要があります。
非同期方式にはデメリットがあります。各ステップのレイテンシを呼び出し全体のレイテンシに合わせて制限する必要があるためです。私たちはこれを極力避けたいと考えており、モデレーションエージェントの非同期バージョンの使用を好んでいます。作成時にはどちらを選んでも構いません。
非同期モデレーションでは、モデレーションエージェントの実行リクエストを受信します。リクエストを受け取ったことを確認した上で、エージェント全体をバックグラウンドで実行します。完了すると、クライアントが購読している Kafka トピックにメッセージを公開し、メッセージが届いた時点でクライアント側で対応できます。
この方式は非常に柔軟性が高く、実行時のレイテンシ制限を大幅に緩和できるため、モデレーションエージェント内でより複雑なステップや追加のステップを組み込むことが可能になります。
モデレーションプラットフォームには、バックテスト機能も備わっています。これは非常に重要なポイントです。プロンプトを作成する際、「これで確実にうまくいくはずだ」と思っても、実際にはそうとは限りません。数例(1〜3 件)でテストしただけでは、本番環境で本当に機能するかを保証できません。
また、先ほど紹介したような複雑なエージェントの場合、どのような出力が得られるかを事前に予測するのは困難です。そこで、プラットフォームのバックテスト機能を活用してエージェントを構築します。過去のデータセットがあれば、そのデータに対してエージェントをテストできます。必要であれば、単一のステップ単位でテストすることも可能です。
人間が各ステップや各エージェントの結果を確認し、「正解」か「不正解」かを分類します。場合によっては、「真陽性」「真陰性」「偽陽性」「偽陰性」といったラベル付けも行うことができます。最終的には、人間の入力に基づいて指標を算出し、本番環境への導入に値する品質かどうか、あるいはさらに微調整が必要かどうかを判断します。
このようにして、「テストしてから信頼する」というワークフローが、当社のプラットフォームには組み込まれています。
Lessons Learned
共有したい教訓が3つあります。
1 つ目は、大量のトラフィックがかかるホットパスで LLM を利用する場合、その前に安価なモデルを配置することです。通常、LLM に到達させるべきではないコンテンツの 90% は、自宅でトレーニングできる安価なモデルでフィルタリングできます。これには時間がかかりますが、自社のデータから学習する必要があります。LLM が行うべき部分は、まず自分で実装し、その上で LLM を活用するのです。このステップを省略してはいけません。誰かが「LLM を導入しましょう」と提案しても、「LLM がないから」という理由だけで安易に採用するのはやめましょう。経済合理性を考えると、LLM に依頼するのは難しい質問に限るべきです。
2 つ目は、LLM にラベル付けや真偽判定(True/False)を求めないことです。代わりに「深刻度」や「スコア」を問うようにしてください。ただし、5 桁のような細かな数値ではなく、LLM が推論しやすいシンプルな形式で回答させることが重要です。
3 つ目は、システムをいつ捨てるべきかを見極めることです。作ったものがうまく機能し、他社も同様のバージョンを求めている場合、そのシステム自体が資産とは限りません。重要なのは「パターン」であり、それを誰もが使える形に変換できる可能性があります。今日ではコード作成コストが非常に安いため、誰でもすぐに実装してしまいます。「再利用する必要はない」という考え方も生まれています。確かに、安価なため自分で作れば済むからです。しかし、「コードを作成するのは安い」ことと、「コードを維持するのは安い」ことは別問題です。
まとめ
本発表では、安全システムを大規模かつリアルタイムなマーケットプレイスで構築する手法について解説します。私たちが開発した低コストのフィルタリング機能は、蓄積されたデータから学習して動作します。また、LLM(大規模言語モデル)を活用した高度な判定エンジンも備えています。これらはすべて、柔軟に拡張可能なプラットフォーム上に統合されています。
このプラットフォームでは、自社インフラでデプロイ・運用する内部モデルと、外部ベンダーが提供する API や契約に基づくサービスを利用できます。特に重要なのが LLM ゲートウェイであり、これはシステム全体の要となるパーツです。
すべての設定は宣言的な構成ファイルで行うため、コードを書き込む必要はありません。UI 上で必要なコンポーネントを選択し、パイプラインを構築するだけで済みます。また、本番環境に投入する前に必ずバックテストを実施することも強く推奨します。プロンプトを盲目的に本番環境に流し込んで結果を確認するのは危険です。
もし現在、自社でモデレーション(コンテンツ審査)に関する課題を抱えている場合でも、このアプローチを採用すれば、比較的容易に解決できるはずです。
質疑応答
参加者1: 内部モデルが「低コスト」だとおっしゃっていましたが、具体的にどの程度の費用感なのか教えていただけますか?内部でモデルを維持・訓練する際にもリソースが必要になるため、ベンダーの低価格帯モデルと比較して、どれほどコストを抑えられているのかイメージしたいです。
Bruna Pereira: コストが安いかどうかは、取引量に依存します。なぜなら「安い」というのは、本質的にサーバーへのデプロイだけで済むからです。これはあなたが持つインフラにも左右されます。重要なのは、呼び出しごとに課金されるのではなく、デプロイする際に一度だけ費用がかかる点です。トレーニング部分には確かにコストがかかりますが、それは一度きりの投資です。例えば SafeChat の場合、すでに 9 番目のバージョンに至っており、9 回トレーニングを行いました。これは 1 日あたり 400 万回の呼び出しに対しての一度きりのコストと捉えています。これが「安さ」や「高価さ」を計算する基準であり、ユースケースやインフラ構成によって大きく異なります。
参加者 2: バックテストの結果を定量化する方法は何かご存知ですか?例えば、「これだけのデータがあれば十分」という判断基準や、さらに多くのシナリオが必要かどうかについてです。
ブルナ・ペレイラ: それは難しい問題です。実際に私たちが直面した「安全性」と「不正行為」の違いについて例を挙げましょう。
安全性の観点では、不適切なメッセージは明確に不適切です。しかし、内部モデルが常に安全かどうかを正確に識別できるわけではありません。
一方、不正行為についてはどうでしょうか。「非常に怪しい内容だが、最終的には問題ないケース」が存在します。例えば、「到着時に支払いをする」という発言や、配達員(Dasher)の親切に対するお礼として追加チップを渡すような場合です。これが本当に不正行為なのかどうかは、そのユースケースがどの程度グレーゾーンにあるかによって判断が分かれます。
通常、私たちは1,000件以上の事例でテストを行うことは望みません。なぜなら、人間が一つひとつ手動で確認し、「真実」か「偽り」かを分類する必要があるからです。もし10万件ものデータがあれば、それを人手で処理するのは不可能です。1,000件という数は、私たちが信頼できる基準となる数値です。これ以上増やすと、担当者たちは疲弊してしまい、適当な回答をしてしまうようになります。その結果、評価スコアも実態を反映しなくなります。
もし数字を挙げるなら、1,000件が適切なラインだと私は考えます。
参加者3: 簡易モデル(cheap model)は、以前から非常に決定論的だったのでしょうか?例えば、単なる悪口のリストのようなものでしたか?それがより複雑化する前には。また、その簡易モデルについてもスコアリングを行うのですか?空港のセキュリティチェックのように、「安全」「不安全」という二択だけでなく、モデル自体の評価も行う必要があるのでしょうか?
ブルナ・ペレイラ: はい、安価なモデルにもスコアリングが適用されます。このスコアは 0 から 1 の範囲で、メッセージの危険度や安全性を示します。
実は、このアプローチにはいくつか課題もありました。最初の学習ラウンドでは、既知の不適切なメッセージを収集してトレーニングしようとしたのですが、問題がありました。なぜなら、私たちが受け取る時点で「すでに十分悪い」メッセージしか集められなかったからです。もしそれほど悪くないメッセージであれば、誰も苦情を寄せないためです。
その結果、最初の学習ラウンドでは「非常に悪いメッセージ」と「まあまあのメッセージ」の二極化が進み、モデルの判断もほぼバイナリ(0 か 1)になってしまいました。
そこで2回目の学習では、市場で利用可能なモデレーション API を活用しました。ただし、この API は非常に遅く、数秒かかるため、私たちの用途には不向きでした。しかし、この API を使うことで、「極めて危険」「やや危険」「安全」といった段階的な評価データを取得できるようになり、モデルからの出力もより漸進的で合理的なものになりました。
結論として、このモデルは単なる真偽(True/False)ではなく、スコアリング機能を持っています。ただし、どのスコアの値で次の層(セカンダリーレイヤー)へ移行するかを定義するのは容易ではありません。例えば、「85% の確率で不適切」と判断された場合、それが即座に次の層へ送られるべきなのかどうかは、ケースバイケースです。
この閾値の調整や学習は、自社のデータに基づいて行い、必要に応じて動的に変更していく必要があります。この点において、ML プラットフォームチーム、データエンジニア、機械学習エンジニアたちは、私たちに多大な貢献をしてくれました。
参加者4:貴社のプラットフォームは、安価で高速なステップなど、他のユースケース向けに新しい内部モデルを構築する自動化も支援しているのでしょうか?
ブルナ・ペレイラ:いいえ、現時点では行っていません。実際、トレーニングや再トレーニング、ファインチューニングの部分は、MLプラットフォームチームが完全に担当しています。私たちはそのプラットフォームに依存して行っており、自前で実施する必要はありません。これは「関心の分離」として優れた設計です。先ほど述べた通り、MLエンジニアが持つ専門知識は私たちが持っていないため、役割を重複させる必要はないからです。当社のモデレーションプラットフォームは、すでにトレーニング済みのモデルを利用するクライアントに過ぎません。
参加者5:先ほど、安価なモデルを数回再トレーニングしていると伺いましたが、そのタイミングや頻度はどのように決定されているのでしょうか?
ブルナ・ペレイラ氏: まず、データに対してランダムな分析を行いました。また、こうした不安全なケースはインシデントとして発生し、エージェントに転送されます。エージェントからは、「このメッセージは顧客に届いたが、モデルで検出されなかった」というフィードバックを受けます。その事実を把握した上で、私たちは見落としたパターンや、それがどのステップで発生したのか(レイヤー1か2か)を探ります。
具体例を挙げましょう。以前、略語に関する再学習を行いました。LLMが略語を理解できるかどうかは問題ではありませんでした。なぜなら、最初のレイヤーで「もちろん安全だ。ここに危険な要素はない」と判断されてしまい、メッセージが先へ進まなかったからです。そこで私は、略語を用いた不安全なメッセージの十分な数の例を見つけ出し、その第一層に対して別のトレーニングセットを実行する必要がありました。
私たちが扱っているデータ量では、通過するすべてのメッセージを人手で確認することはできません。通常はフローが安全担当者に到達します。彼らから報告を受け、私たちは「見落としたパターンは何だったのか」を特定しようと努めています。
参加者5番: プラットフォームのスマートな部分に着手する際、外部モデルと接続されますね。もし外部モデルが遅くなりすぎた場合(メッセージには低レイテンシが求められます)、フォールバックはどうなりますか?
ブルナ・ペレイラ氏: 実は、ここで言及しなかった点があります。フォールバック(代替手段)にはあらゆる種類のモデルを採用可能です。当社には内部で開発したモデルがありますが、サーバーが障害を起こした場合に備え、外部ベンダーによるモデレーション層をフォールバックとして用意しています。サーバーの復旧までの間はこれを利用します。速度は劣りますが、何もしないよりはマシです。
このフォールバックには何でも使えます。例えば LLM(大規模言語モデル)を採用することも可能です。ただし、チャット専用としては使いません。処理ボリュームが大きすぎるためです。フォールバックとしてあらゆる種類のモデルを追加できます。特定のベンダーが正常に動作していない場合、API の復旧中は LLM を一時的に利用し、復旧後に再びそのベンダーに戻すといった運用も可能です。
参加者 6 氏: バックテストについて少しお話しください。人間の介入が必要になる点に触れられていましたが、そのデータはプロジェクト開始前に既に存在していたものですか?それとも、プロジェクトを始めた際に新たに計算・作成したデータでしょうか?すでに不審なメッセージのセットが用意されていたのでしょうか?あるいは、過去のデータを精査して該当箇所を見つけるプロセスの一部として準備されたのでしょうか?
Bruna Pereira: 当初、チャット機能が計測されておらず、データが極めて限られていました。学習用のデータもありませんでした。そこでまず着手したのは、言葉による虐待によって作成された安全対策事例を調査し、そこから不適切なメッセージの具体例を抽出することです。その後、チャットの計測を開始しました。初期段階では、データの収集に約1年半を要しましたが、この時点で必要なデータを揃えることができました。これにより、必要に応じてモデルを再学習させる体制が整いました。
しかし、いくつか見落としもありました。例えば、略語の検出は当初できていませんでした。すべてのメッセージを確認しているわけではないためです。また、人間からのフィードバックを通じて新たな課題を発見することもあります。「ここに問題がある」と指摘された事例からパターンを特定し、プロンプトや既存モデルに組み込むことで改善を図ります。
具体的な例として、当初は一定のカテゴリ分類を設定していましたが、人間が対応するケースの多くが「安全性」の問題ではなく、「不敬」に関連していることに気づきました。そこで、LLM に「不敬度合い」を判定するためのカテゴリを追加し、メッセージの不敬度を評価できるようにしました。
一度設定すればそれで終わりではありません。継続的なループと改善が不可欠です。
参加者7: 詐欺というより微妙な問題の解決に取り組み始めた際、プロンプトインジェクションやその他の新たな課題にも対応する必要が生じましたか?
ブルナ・ペレイラ: この問題については、すでに適切な対策を講じていたと考えています。特に LLM ゲートウェイ(LLM Gateway)は、このようなコンテンツのブロックや制限に役立っています。確かに、不正利用のケース以前から、そのような仕組みは用意されていました。
通訳付きプレゼンテーション をもっと見る
録画日:
2026 年 8 月 22 日
原文を表示
SafeChat: Building AI-Powered Safety Systems at Scale in a Real-Time Marketplace
View Presentation
Speed:
Download
42:22
/presentations/doordash-llm-ai-moderation-platform/en/slides/slide-1786348150151.jpg)
まとめ
Bruna Pereira explains how DoorDash built a content-agnostic AI moderation platform. She covers replacing costly LLM-only pipelines with a hybrid pattern: using fast internal models to filter obvious cases, LLM multi-axis scoring for nuanced decisions, and no-code workflows with backtesting. Discover how this architectural pattern cut safety incidents while scaling to millions of daily messages.
Bio
Bruna Pereira is a software engineer at DoorDash with 10+ years of experience in software engineering. She enjoys solving hard problems, working on systems that scale, and learning from building things in production.
About the conference
QCon AI is a practitioner-led event focused entirely on the engineering discipline required to scale these workloads safely. It provides direct
access to the architectural playbooks and failure metrics that peer organizations use in production.
INFOQ EVENTS
- August 27th, 2026, 1 PM EDT
Why AI Agents Fail at Scale: Context Is a Data Layer Problem
Presented by: Boyd Stowe - Founding Solutions Architect at Tacnode
- September 17th, 2026, 1 PM EDT
Beyond the PR: The New Control Plane for Agentic Software Delivery
Presented by: Mohit Suman - Staff Product Manager at Harness
Transcript
Bruna Pereira: I'm Bruna. I'm a software engineer at DoorDash. Today, I will take you on a story about a system that we built to solve a real safety problem in production, and when it worked, we threw it all away to build something even more powerful. By the end, I want to leave you with two things, how we built SafeChat at DoorDash, and one architectural pattern that you can take home and use at almost any AI use case that you have. I've been a software engineer at DoorDash for one-and-a-half years now. I lead the trust and safety engineering team from our hub in Sao Paulo, Brazil. Before that, I spent three years building my own startup. My background is mostly on FinTech, building highly scalable systems.
The DoorDash Marketplace (Safety-First)
DoorDash is a marketplace. We have an app that consumers can order food from. We have the Dashers that deliver the food. We have the merchants that prepare the food. These three groups, they can meet in different parts of the journey. They can talk to each other in chat. They can call each other during active deliveries, and of course they can meet in person. For DoorDash, two things matter equally for us, that people are safe and that people feel safe in the platform. For us, both metrics are equal product metrics for us. When we talk about safety, a meaningful amount of the safety incidents that we see on the platform are related to verbal abuse. These groups, since they meet in chat or in voice, it doesn't matter. We need to act right away when we sense that there's something unsafe happening, especially because different from a social media platform, we don't get to build strong relationships between these parties.
These relationships, they last from 40 minutes to 6 minutes max. We must act right away when we see something unsafe happening. Talking about scale, we have over 4 million messages exchanged in chat every day. Over 400,000 calls are exchanged between consumers and Dashers during the deliveries. Over 200,000 images are exchanged in chat or even SMS. For chat specifically, our goal is to make sure that every message that reaches the destination is classified as safe by us. This is what leverages the difficulty of this system, because look at the amount of messages that we have. For chat, we only have a fraction of a second to decide if this message is safe or not. Otherwise, we would be disturbing the chat experience. When we first talked about it, the first response, especially from business folks, was to just throw an LLM on the problem. Just call an LLM, ask if the message is safe or not.
If it's safe, deliver it. If it's not, block it. Yes, it would have worked theoretically, but in practice with the volume that we have, if you have used LLMs before in production context, you should know that latency of these LLM calls, they can vary a lot. For our use case, we have like 2 to 10 seconds in average of latency for each of these calls. Also, imagine what is the size of the bill of calling the LLM 4 million times a day. It would just not work.
Building SafeChat
The question became, how do we be smart about the usage of LLMs? The first thing that we did is we tried to understand our data. We spent a couple of months understanding what unsafety means in our context, and also, what is the percentage of messages that is actually unsafe. Because we know that most messages are unsafe, but how many messages? The point was not to start making decisions here. The point was just to learn from the data that we have. We instrumented the chat. We used a free moderation API available on the market. We asynchronously called this API to understand, what's the category of unsafe messages that we have here? This is the boring step that everyone wants to skip. This is the conversation that I have a lot back at work. Let's spend a couple of months understanding our data before we act, because it doesn't matter how cheap or expensive the model is.
If you don't know your data, you won't be able to get anything out of it. When we ran this analysis, we confirmed what we already knew. Most of the messages are safe, but we actually put a number on it. Only a small single-digit percent of messages were unsafe. This shaped our entire architecture, because it told us that we could build an aggressive cheap layer that was right almost all of the time.
Then we built a small classifier. It is an ML model trained on the data that we collected in the previous step. It had three jobs. It had to be fast and respond in less than 100 milliseconds at 90%. It had to be cheap, so no per-call costs, and only the cost of our infrastructure. It had to be good at one thing, identifying what is obviously safe. This layer is not a final judge. This layer is, just imagine that you are at the U.S. airport and you have the metal detector. A metal detector does not say that you are threatened. It just says that you are worth a closer look. For the messages that this model that we trained in-house is not able to classify as safe, then we call an LLM, which is less than 10% of the messages. Here's the design choice that matters.
When we call the LLM, we don't ask it if the message is safe or not. Instead, we ask it to classify or to score across multiple axes. I will tell you in a couple of slides why this is important. What we ask is depending on the context, in our context, we ask how threatening this message is, or how profane this message is, or how sexual this message is. This is the layer where the model gets to be smart, because you only ask it the hard questions.
Talking about the overall architecture here, a message comes in and we strip out all of the noise. We remove the empty messages, the image attachments, because those go to a different pipeline. We remove the common pleasantries-like things. Then it goes to this first layer that is the small model, the classifier. If the message is safe, we just ship it. If it's not, it goes to the layer 2 which is the LLM. The LLM classifies across multiple axes. With that result, we can take a graduated action. This is the shape that I'll show you through the entire talk. The cheap layer that goes most of the messages, the expensive layer, and then action. Each layer is doing the job it's good at, at the right cost. That's the important part. This is the slide that I want you to remember if you ever build your moderation system.
Do not ask the LLM for a Boolean. Instead, ask the LLM for a score. The reason for that is because a Boolean is a flag and a score is a knob. With a score, you can take graduated actions depending on the score that you received. You can add new categories later without having to go back and recreate everything. You can move the thresholds if you think that it's too high or too low. Also, it plays at what LLMs are good at. Because imagine if I throw a sentence here to you and say, is this message safe or unsafe? Probably your answers will be all different, because what is unsafe in your context and how would we classify something as unsafe? Instead, understand what unsafety means and ask the LLM to score across this axis. The results will be way more stable.
This is why the score is important as well, because we must act according to the severity of the message. For us in our context, low-severity content means like a swearing. If you swear, we just censor the message and let it go through. If it's a mid-severity message, for example, you insulted someone, we can block the message entirely and not let it go through. If it's a high-severity message, for example a threaten, we can block the message and offer the affected party to cancel the order without paying anything. If it's a very high severity content, we just also cancel the order, block the message, warn the offender, and remove the affected party completely from this loop. This letter is only possible because we have those scores. If we didn't have it, we wouldn't be able to differentiate what is rude and what is actually dangerous. I talked about chat, but for voice and image, the structure is pretty much the same.
We have the same scoring engine underneath. Some things are different. For image, for example, instead of that cheap internal layer, we have a commercial vision API that is already good enough at identifying violence, nudity, and unsafe images. We use that as the cheap layer. For voice, it's tricky, because for chat, we can read the message and decide if we ship it or not. For voice, at the moment that we transcribed the message and analyzed it, the words were already heard by the recipient. There's no way that we can avoid the message from being delivered. The only thing that we can do differently is to act. We can hang up the call as soon as we identify that there's something unsafe there, and we can, of course, offer the order to be canceled or cancel the order ourselves. The engine is pretty much the same.
The Result
After we did that, we measured how many incidents we had driven by verbal abuse before and after we implemented what we call SafeChat. We saw that we had roughly a 50% reduction in those incidents. This number is not like a model accuracy improvement that we had. This is real reduction in human harm. This is the number that we can use to justify the months that we spent understanding our data to build something powerful enough to help our customers.
A Content-Agnostic Moderation Platform - Building Blocks
This is the moment that I would finalize the talk and ship the product, and the team would get promoted and everyone goes home. What happened was that we threw it all away. Not the learnings that we had. Not the model that we trained. Not the data that we collected. What we threw away was the system, the SafeChat system. I'll tell you why we did that. When we started rolling out SafeChat, other people at DoorDash and from other teams, they were like, "I like what you're doing. I want to do it myself as well for my use case." People started asking us, can we moderate profile pictures for Dashers and consumers? Can we moderate name at signup? Can we moderate food reviews, for example? Even for cases that are not related to safety at all, like can we identify fraud in chat and in phone calls?
We noticed that if we started to implement each of these asks, we would be rebuilding everything over and over from scratch. We noticed that what they actually wanted was not SafeChat, was not the system that we built. What they wanted was the pattern that we created. That cheap, then expensive, then graduated action. That is what they wanted. We thought, why don't we transform it into a platform that anyone can use for any use case? That's what we built. We built a content-agnostic moderation platform. The idea here is that we don't need to know what your content means. We don't need to know your business logic. What we know, what to do, is how to log decisions, how to integrate with different model providers, how to change steps with conditions between them. The teams, they bring the meaning and the platform orchestrates that. That's the idea.
What we promised to the teams is that they can do that without writing any code. They can go to the platform in a UI. They can configure it and they can take the action that they want. That's the flexibility story here. We did that. It doesn't matter if you have a content guideline or a fraud use case or a safety use case. It doesn't matter for us. The platform, it has some building blocks that make it possible, and I will show you what they are.
We have three kinds of models inside of this platform as of now. We have what we call internal models that are models that we can train, fine-tune, and deploy in our own servers. We have external models, which are models that we can just sign a contract and use it from a vendor. We have an external prompt. An external prompt is a prompt that you can write to use in any LLM from any vendor that we have on the market. For internal models, they are trained and hosted internally. We have an ML platform that helps us to use any model available on the market, use our own labeled data. If we have labeled data, we can use that to train these models. We can deploy that in our own infrastructure. These models, they are served under an API that we can use inside our platform. The idea is that they are cheap and fast.
When we are training it, we have a predefined input and output schema that when we defined it, everyone, every client can use following this input and output schema. The small classifier that we built for SafeChat, it's actually an internal model now. We have the external models. The idea here is not to reinvent anything that is already available on the market. For example, image safety is something that is already solved. We don't need to build it all over again. The idea here is that the platform does not rebuild stuff just because. We have it available. We have a contract with them. We integrate once in our platform and any client is free to use. We just have different API keys so we can trace billing accordingly. We have the external prompts. The external prompts are the most flexible ones. We use an LLM gateway to integrate with pretty much any model from any vendor available on the market.
If you are more interested in knowing how we built this LLM gateway, there is a talk from DoorDash folks about how we built the LLM gateway at DoorDash. The idea here is that this gateway sits between the platform and all the models from all providers available outside. The clients can choose what models they are targeting. They can choose the input schemas and the output schemas for each prompt that they have. We can use the features that are available in this LLM gateway to declare fallback and retry strategies without writing any code again. We use the fallback and retry to, for example, if you've already used LLM models, you must know that sometimes they just are not responding. The providers are all down or the models are not working as expected. If you have created a prompt that targets a model that is not available, is there any other prompt that you can use as a fallback while this is not back healthy? Also, you define what is the output schema that you want from that LLM call. You define that JSON. Sometimes the LLMs, they just don't respond to that. You can define how many times you want to retry before returning a structured error to the client.
Composing Moderation Agents
Those are the building blocks. With these building blocks, you can compose into what we call moderation agents. A moderation agent is a pipeline or a workflow that you can compose the steps. Here is an example of an internal model that you can choose a condition. This condition is going to use whatever is returned from this internal model. With this condition, you can go to one step or another. In this case, it goes to an external vendor or an LLM prompt. You can change it, however, and you can take an action in the end. The action is the responsibility of the client. The client decides. At this point, the client has the output of each step that they have in their moderation agent. With that, they can take any action that they have, or no if needed. Talking about condition, the way that we express conditions is, again, entirely from the UI.
You can write and express any conditions that you want, simple or complex, using the output from the previous model. This is actually the pipeline from SafeChat. It's super simple. It has an internal model. If the unsafe label that is returned from this internal model is greater than 0.5, then it goes to the LLM prompt. The LLM prompt is going to respond with a set of categories and the score of each of them. It goes to an action. If this is not greater than 0.5, it just goes to an action. Here I know that we don't take any action. Actually, we censor the message if there is any swearing in it. You can express any condition between any steps and you can create your moderation agent. Your moderation agent, it can be of two different types. It can be a synchronous moderation agent, which means that it's an HTTP call.
The connection is held open while the moderation agent is being executed. We execute all the steps that we have in the agent. We use synchronous only when we want to gate a decision. For example, in chat, we want to be able to block the message if the message is unsafe. There is a downside of using asynchronous because you need to cap the latency of each of your steps according to the overall latency of your call. We try to avoid that as much as possible, and we prefer to use the async version of the moderation agent. You can choose whatever when you are creating. For async moderation, we receive a request of moderation agent execution. We acknowledge that you asked that, and we run the entire agent in the background. When we finish it, we publish a message to a Kafka topic that the client is subscribed to, and they can react to it when the message comes. It's way more flexible. It helps you to not cap so much the latency of the execution, and you can add more complex steps and more steps actually inside the moderation agent.
One other feature that we have on the moderation platform is the backtesting. This is interesting because sometimes when we create a prompt we are like, that's going to work for sure but we are not that sure. We test it against one, two, three examples but it's not enough to make sure that it actually works in production. Also imagine an agent as complex as I showed before, you probably don't know what's going to come out of it. The idea is that you can use the backtesting functionality from the platform to build your agent. If you have some historical set of data, you can test your agent against this set of data, and you can even test it against a single step if you want to. A human can go there and classify the results of each step or each agent as correct or incorrect, and in some cases we can even label it as true positive, true negative, false positive, false negative. By the end we use that information that the human input to calculate some metrics to understand if it's good enough to go to production or if we need to fine-tune it even harder. This is what makes the test it before trust it a real built-in workflow in our platform.
Lessons Learned
I have three lessons here that I want to share with you. One of them is, if you have a hot path with high volume and you want to use an LLM for it, put a cheap model in front of it. Usually, the 90% of content that you are sure that should not reach the LLM, it can be caught by a cheap model that you can train at home. It takes some time. You need to learn from your data. That part that the LLM does for you, you need to do it yourself first. Don't skip that part. Resist that. If someone asks you, yes, but let's add the LLM just because we don't have, no way. Because the economics, they only work if you ask only the hard questions to the LLM. When using LLM, do not ask for labels. Do not ask for, is it true or false?
Ask it for a severity, ask it for a score. Also don't ask it for a 5-digit score because the LLMs are not calculating anything, they are reasoning. Ask it for a score in a simple way that the LLM can answer that. Know when to throw a system away. If you build something and you see that it's good, and people are asking for their version of it, maybe the system that you built is not the asset that you have. Maybe the pattern is, and you can transform it into something that everyone can use. I noticed that today it's so cheap to create code that people just go for it. Yes, we don't need to reuse because we can build our own because it's that cheap. It's cheap to create code, it's not cheap to maintain the code.
まとめ
This is the whole talk in one shape. We have this cheap filter that we built, learning the data that we have. We have the smart judge using LLMs. We have the graduated action. All of that on top of a platform that we can use. Internal models, models that we deploy and serve from our own infrastructure. We have vendors, contracts, APIs that we can use and provide to any client. We have the LLMs. The LLM gateway is actually a very important piece of this platform. All of that with declarative configuration, so no need to code for that. It's already set in a way that you just need to select the pieces in the UI and create your own pipeline. The backtesting is super important as well. It's terrible to go blindly with some prompt to production to see what happens when you throw it there. I'm pretty sure that if you have a moderation problem at home, if you use that pattern, you'll be able to solve it in an easy way.
Questions and Answers
Participant 1: You mentioned that the internal models were cheap. Can you give us a sense of how cheap? Because I imagine there's resources that goes into maintaining and training those models internally too. How cheap is it compared to lower tier models from the vendor itself?
Bruna Pereira: How cheap it is depends on the volume that you have. Because the cheap is in a sense that we only need to deploy it to our servers. It depends on the infrastructure that you have. The point is that you do not pay per call. You pay just to deploy it. The training part, yes, that has some costs that you do it only once. Actually, for the SafeChat, for example, we are in the 9th version. We trained it nine times. It's like a one-time cost for 4 million calls a day. That's how we calculate cheap and expensive. It depends a lot on the use case and on your infrastructure.
Participant 2: Did you come up with any way to quantify the backtesting, like how much is enough or we need more kind of scenario?
Bruna Pereira: That's hard. I'll give you an example that we just used, that is safety versus fraud. For safety, the unsafe messages, they are clearly unsafe. The internal model cannot always identify what is safe or not. For fraud, for example, there are things that are pretty fraud-y, but then they are not in the end. Because they say, yes, I'm going to pay you when you get here. They want just to give some extra tips to the Dasher because the Dasher is doing a favor. Is it fraud or not? It depends on how much of a gray area your use case is. Usually, what we do is we do not want to like test over a thousand examples, because we need a human to go over that manually and classify what is true and what is not true. If you have 100,000, you won't be able to do it. A thousand is a number that we actually trust. People stop to do that, and care about that. Because we tried more and people would just answer with anything, so it reflects in the scores that we have. If I had to throw a number, I would say a thousand is a good number.
Participant 3: Was the cheap model ever super deterministic, like just an array of swear words before it got more complicated than that? Then, do you also score cheap model instead of just having an airport security, yes, no, this is safe, this is not safe?
Bruna Pereira: Yes, the cheap model is scored also. It has like a 0 to 1, how unsafe and how safe it is. Yes, we did have some problems with that, because when we ran our first round, we tried to just collect some unsafe messages that we had. When the message comes to us, it's already bad enough. Because if it's like, not that bad, no one is going to complain. The first round of training that we had, it was like a bunch of really bad messages and a bunch of ok messages. The model was a bit of almost binary. Then we ran a second round of it, we used this moderation API that we have available on the market, which is super slow. That's why it does not work for us. It takes seconds to return something. Then it could help us with data that is, yes, this is super unsafe, this is kind of unsafe, and this is safe.
It helped us to return gradual and more reasonable results from the model. Definitely the model is not a true and false. It's also a score. It's hard also to define what is the score that you say that with this score, you should go to the second layer. Because if it says yes, then 85% chance of it being an unsafe message, does it go to a second layer or not? This is something that you need to learn from your data and move the thresholds when you need. With this, the ML platform team, the data engineers, the machine learning engineers, actually, they helped us a lot with that.
Participant 4: Does your platform also help with automating the building of new internal models to fit other use cases like for the cheap and fast step?
Bruna Pereira: No, not right now. Actually, the training and retraining and fine-tuning part is something that is totally handled by the ML platform team. We rely on their platform to do that. We do not need to do it ourselves. This is a good separation of concern because there is totally knowledge, as I mentioned here, that the ML engineers they have and we don't. We do not want to have this overlap. The moderation platform is only a client of this model that is already trained.
Participant 5: You mentioned that you retrain the model a couple times, the cheap one. How do you decide?
Bruna Pereira: First, we did random analysis on the data. Also, these unsafe cases, they create incidents and they go to agents, and the agents feed us back with the information of, yes, this message came to the customer and it was not caught by the model. Then when we note that, we try to understand what's the pattern that we missed and in which step? It was in layer 1 or 2? I'll give you an example. We had a retraining. The last time that I retrained, I retrained it on abbreviations. It doesn't matter if the LLM understands the abbreviation because it never reached there because the first layer says, of course, it's safe. There is nothing unsafe here. I had to find a meaningful amount of examples of unsafe messages that use abbreviations and then run another set of training in this first layer. With the volume that we have, we cannot see all the messages that go through. Usually, the flow reaches a safety person. They come to us and we try to identify what is this pattern that we missed.
Participant 5: When you go for the smart piece of the platform, you're connecting the external model. What's the fallback if this becomes too slow, for instance, because you mentioned that you need low latency on the messages? What's the fallback?
Bruna Pereira: Actually, one thing that I didn't mention here is that the fallback can be any kind of model. We do have this internal model, but sometimes our servers fail. We have a fallback with this moderation layer that is an external vendor. We use it while we don't recover our servers. It's slower, but it's better than nothing. The fallback here can be any. You can use, as a fallback, the LLM, for example. We wouldn't use for chat specifically because the volume is too high. You can add as a fallback, any sort of model. If you have a vendor that is not working properly, maybe you can use an LLM just while the vendor is recovering their API, and then you can return to use the vendor.
Participant 6: Just talking a little bit about the backtesting, you talked about a little bit the human need to actually produce that. Is that data that already existed or is that data that had to be calculated as you began this project? Did you have a whole set of unsafe messages already established or is that something you had to review old data and find those as part of this process?
Bruna Pereira: When we started, we had only a few data because we didn't have it instrumented, the chat. We didn't have any data to train on. What we did first is that we found the safety cases that were created due to verbal abuse. From that, we could find examples of unsafe messages. We instrumented the chat. When we started, we spent like one-and-a-half months only collecting data. From that time on, we had the data that we needed to retrain it whenever we wanted. Some things we missed. For example, the abbreviation part we missed because we don't get to look at all the messages. Sometimes we learn stuff when a human comes to us and says that something is wrong, and then we can find the pattern and integrate that pattern in the prompts or in the models that we have. One example that we had is like, we started with some set of categories and then we realized that most of the complaints that reached a human were related to disrespect and not, that's unsafety. We added a disrespect category in our LLM to identify how disrespectful this message is, for example. Yes, you don't do it once and then it works forever. You need to keep looping and improving it.
Participant 7: As you started solving the more nuanced problem of fraud, did you also find you had to start tackling prompt injection or other emergent problems?
Bruna Pereira: I think we were already smart about this problem. Especially the LLM gateway, it also helps us with blocking or not allowing this kind of stuff. Yes, we do have some ways to handle that. Even before fraud use cases, we already had stuff like that, definitely.
See more presentations with transcripts
Recorded at:
Aug 22, 2026
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み