Anthropic クロードコードチームが語る AI エンジニアリング
Anthropic の Claude Code チームは、開発プロセスにおける AI エージェントの活用率やシステムプロンプト設計の最適化に関する具体的な内部データと戦略を公開し、業界標準の見直しを促した。
キーポイント
Claude Tag の高い採用率と機能実装
Anthropic 内の製品エンジニアリング PR の 65% が Claude Tag(Slack 統合型 AI エージェント)によって処理されており、新機能はまず社内でテストされ、ユーザー定着率が確認されたもののみが一般公開される。
システムプロンプト設計のパラダイムシフト
最新のモデル(Fable 5 や Opus 4.8)では、例示や「〜してはいけない」という禁止事項を列挙する従来手法が非推奨となり、Claude Code のシステムプロンプトは 80% 削減された。
自動レビューと手動レビューの役割分担
クリティカルな変更点は依然として手動レビューだが、製品の「外側」の層については自動化されたコードレビューへの依存度が高まっており、開発効率化が進んでいる。
社内文化と新機能の実証
社内外での公開的なコラボレーション(Ant Fooding)が成功の鍵であり、Fable モデルは動画編集にも有能で、同チームが自身のローンチビデオ制作に使用したことが示された。
重要な引用
Claude Tag (Claude's new collaborative Slack integration) now lands 65% of the product engineering PRs for the Claude Code team.
Adding examples to a system prompt is no longer best practice for models like Fable 5 or even Opus 4.8.
The Claude Code system prompt recently reduced in size by 80%.
Lists of 'don't do X and don't do Y' can reduce the quality of results from the latest models.
影響分析・編集コメントを表示
影響分析
この記事は、大規模言語モデルの進化が単なるチャットボットの性能向上にとどまらず、開発ワークフローそのものを再定義していることを示唆しています。特に「禁止事項を減らす」「例示を減らす」という逆説的なプロンプト設計の変化は、AI モデルの推論能力が飛躍的に向上した証左であり、開発者が従来のプロンプトエンジニアリング手法を見直すきっかけとなるでしょう。また、社内の高い採用率と公開文化は、他社における AI エージェント導入の成功事例として重要な示唆を与えます。
編集コメント
Claude Code チームの内部データは、AI エージェントが単なる実験段階から実務の核へと移行したことを如実に示しています。特にプロンプト設計における「禁止事項の排除」や「例示の削減」という逆説的な知見は、モデルの推論能力が飛躍的に向上した証左であり、開発者が従来の手法を見直す重要な契機となるでしょう。
今月初め、私は「AI Engineer World's Fair」(2026 年開催) で、Anthropic の Claude Code チームに所属する Cat Wu 氏と Thariq Shihipar 氏を招き、フアイアサイド・チャットを開催しました。議論のテーマは、Claude Code や Claude Tag、Fable といったツール、コーディングエージェントのセキュリティ、評価手法 (evals)、ツールの設計思想、そして Anthropic 自身がこれらのツールをどう活用しているかについてでした。
セッションの完全な動画は現在 YouTube で公開されています。以下に、追加リンクや私の強調を加えた編集版 transcript を掲載します。
*
もし動画視聴や全文の読み込みが難しい場合のために、主要なポイントをいくつかまとめました:
Claude Tag(Claude の新しい共同作業用 Slack 統合機能)は、現在 Claude Code チームの製品エンジニアリングにおける PR の 65% をカバーするまでに成長しました。
Claude Code では新機能がまず Anthropic 社員向けにリリースされ、そのユーザー層での定着率が確認されたものだけが一般公開されます。重要な変更点は依然として手動レビューの対象ですが、製品の「外側」のレイヤーについては自動化されたコードレビューへの依存度が高まっています。
Fable 5 や Opus 4.8 のような最新モデルにおいて、システムプロンプトに例を追加することが最良の方法ではなくなりました。Claude Code のシステムプロンプトは最近、サイズが 80% 削減されています。同様に、「X をしない」「Y をしない」といった禁止事項のリストを列挙することも、最新のモデルにおいては結果の質を下げる要因となり得ます。
Anthropic 社内での自社製品活用は「ant fooding(アリ食い)」と呼ばれています。Anthropic は自動モードに強い信頼を寄せており、これを Claude Tag の基盤技術と捉えています。
Thariq は、コーディングエージェントによる深い没入感(Deep Blue)に対抗するためには、「取り組む仕事に対してより野心的であること」が重要だとアドバイスしています。
Fable は動画編集にも対応しており、Thariq 自身が Fable を使って製品発表用の動画を編集しました。
社内での作業を公開する文化こそが Anthropic の成功の鍵であり、その象徴として公共の Slack チャンネルで Claude Tag を活用している様子が示されています。
Simon: クロード・コードは昨年 2 月に登場しました。まだ 1 年半も経っていない新しいツールですが、当初は Claude Sonnet 3.7 の発表の項目の一つに過ぎませんでした。実際に私たちをサポートするコーディングエージェントが誕生した今、日々の業務はどのように変化しましたか?
Cat: クロード・コードと Sonnet 3.7 を初めてリリースした頃を思い出します。その頃はタスクを与えると、一つ一つの動作を細かく監視する必要がありました。私はすべての権限プロンプトに非常に慎重に対応し、「いいえ」と答えることも頻繁でした。「このファイルを確認しましたか?」「あのファイルも確認しましたか?」と。しかし、モデルのバージョンが上がるにつれて劇的に改善されました。私たちは皆、些末な実装作業をクロード・コードに任せることで一歩引く機会を得ました。これにより、クリエイティブな仕事に集中する時間が大幅に増えました。「クロード・コードなら多くの機能を実装できる」という前提のもと、「ユーザーにどのような体験を提供すべきか」を考える時間です。そして Fable の登場は、さらに一段階大きな飛躍をもたらしました。Fable を使えば、多くのユースケースで一度の指示だけで多数の機能を構築できるようになりました。
Thariq: Claude Code に関する最初の連絡を覚えています。親友の一人から「Claude Code を試す必要があるよ」と言われたんです。Opus 4 がリリースされた頃の話で、実際に使ってみて「おっと、これはヤバい。Anthropic で働くべきだ」と思ったものです。あの Opus 4 は素晴らしいモデルでしたが、それでも許可プロンプトを読み込む必要がありました。
自動モードが昔から存在していたかのような錯覚に陥るほど、私たちは記憶を失っているのかもしれませんね。自分が「はい」や「許可」を押したことをさえ覚えていないのです。私自身、今特に力を入れているのは、これまでにない高品質な仕事をする必要があるという点です。出力の質は非常に高いですが、私は動画編集に頻繁に使っています。ブランドチームが求める極めて厳格な基準を数時間でクリアできなければ、プロジェクトは成立しないのです。
これが私が Fable で目指している変化です。これまでで最良の仕事を、これまで以上に速く成し遂げることです。
従来のソフトウェアエンジニアリングのどの常識がもはや通用しないのか?
Simon: 一年前までは真実だったとされる、従来のソフトウェアエンジニアリングにおける「常識」で、新しい世界ではもはや通用しないものはありますか?
Cat: 私たちが今、エンジニアリングのスキルセットにおいて最も大きな変化として感じていることのひとつは、2 年前と現在では状況が逆転している点です。かつては、プロダクトマネージャーが顧客に直接話を聞き、6 ヶ月かけて横断的なチームと合意形成を図り、コードを一行も書かれる前に「どのように実装するか」を詳細に記した仕様書を完成させるのが一般的でした。
しかし現在は状況が全く逆です。私がこの場にいるエンジニアの皆さんにお勧めするのは、私たちが何を構築すべきかについて、ビジネス感覚やプロダクト感覚をさらに磨いてほしいということです。アイデアから実装までのタイムラインが劇的に短縮されたからです。6 ヶ月〜12 ヶ月かかっていたものが、今ではわずか 1 週間程度で済むケースもあります。
つまり、私たちが何を構築する価値があるのか、実際に事業にどのような影響を与えるのかについて、より優れた「味覚(センス)」を持つ必要があるのです。多くのプロダクト領域では、「実行力」よりも「プロダクトの味覚とビジネス感覚」の重要性が増していると言えます。もちろん、インフラ分野では、細部まで正確であることの重視という点は依然として変わりません。
Thariq: 私にとって重要なのは、リワーク(書き直し)がもはや悪いことではないということです。
Simon: かつては最悪の行為とされていたことが、今や許容されるようになっているのです。
Thariq: その通りです。『ミステリアス・マン・マンス』の教訓にある「書き換えはするな」という考え方は、今は否定します。むしろ私は書き換えを支持しています。良いテストスイートがあればこそ——実は書き直しを行うことで、良いテストスイートを整える必要性に気づかされるのだと思います。ただ、人々が過小評価しているのは、コードベースそのものが仕様書であるという点です。場合によっては、それが唯一の仕様書の写しになることもあります。なぜなら、誰もコードベースのすべての分岐を完全に把握しているわけではないからです。このコードベースを一つの成果物として捉え、要約したり、別のバージョンを作成したりすることも可能です。実際にBun を Rust で書き直した事例がありますが、これは非常にうまく機能しています——私自身も現在、これを実環境で利用しています。
Simon: Claude Code はまだ Bun in Rust 版としてリリースされていないんですよね?
Thariq: 社内では既に展開済みです。
(実際には Anthropic は、6 月 17 日付で一般ユーザー向けに Claude Code の Bun in Rust 版の提供を開始したようです。)
エンジニア以外の人は Claude Tag で何をしているのか?
Simon: この他にも最近大きな話題となったのが、Claude Tag です。少なくとも一般ユーザーにとっては、すでに一週間以上が経過しています。Anthropic 社内ではエンジニア以外の社員も Claude Tag を非常に活用しているとのことです。具体的に、エンジニア以外の人は Claude Tag でどのようなことをしているのでしょうか?
Cat: Claude Tag は、チームの協働ツールに常駐する Claude です。先週、Slack 内で公開しました。Claude Tag の最大の特徴は、デフォルトでマルチプレイヤー対応している点です。Slack チャンネルに Claude Tag を追加すれば、あなたも参加でき、同僚も参加できます。そして共同でプルリクエスト(PR)に取り組むことができます。
もう一つ大きな違いは、受動的ではなく能動的であることです。「このチャンネルのすべてのバグ報告を監視し、修正用の PR を作成して、そのコード部分を最後に更新したエンジニアにタグ付けして」と指示すれば、チャンネルが存続する限り、毎回手動で呼び出すことなく自動的に実行してくれます。
3 つ目の大きな変化は、チームメモリ機能を追加した点です。チャンネル内で Claude Tag にあなたの好みを伝えれば、今後の投稿でもそれを記憶します。「障害のデバッグはしてほしいが、警告のデバッグは不要」といった設定を自然言語でチャンネル内に伝えるだけで、あなたとチームメンバー全員のためにその設定を覚えておいてくれます。
社内の見解では、Claude Tag は Claude Code の進化形です。これは社内での働き方における大きな転換点だと捉えています。現在、Claude Tag が処理するプロダクトエンジニアリングの PR 比率は 65% に達しています。
Simon: それは Anthropic 全体の話ですか?それとも Claude Code だけの話ですか?
Cat: これは製品エンジニアリングチーム向けの話ですが、社内で使っているバージョンの Claude Tag は、現在、製品のプルリクエストの 65% を処理しています。これは大きな転換点で、全 PR の半数以上を占めるようになりました。Claude Code と Claude Tag の役割分担については、以下のように捉えています。最も複雑なタスクや、エージェントと対話しながら反復して作業する際には、依然として Claude Code が最適です。一方で、Claude Tag は「自分たちの代わりに積極的に作業させる」のに適しています。これにより、開発中の機能で発生するバグ報告に対して、いちいち手動で Claude Code を起動する必要がなくなります。
Thariq: コーディング以外のケースでも活用できます。例えば今回のトークの前、私たちは Claude Tag に「Fable のリリース日はいつですか?」と尋ねました。発表に合わせて準備を整えるためです。Claude Tag は社内の Slack を検索し、誰が何を発言したかを確認します。これは企業内における検索エンジンとして非常に価値があります。製品に関する文脈をすべて把握しているため、数値に関する質問にも答えることができます。意思決定をする際、その根拠となるデータを提示してほしい場合が多いからです。そのため、イベントストアに接続して活用しています。マーケティングチームでは、「この機能について教えてください」といった使い方もしています。彼らはプログラマーではありませんが、Claude はプログラマーとして振る舞い、コードベースをクローンした上で「これがその機能で、こんな見た目です私が実際に使っている様子を録画しました」と説明できます。これにより非常に多様な活用が可能になり、私たちはまだその可能性を探り始めたばかりだと考えています。
チームでの協働を担う Claude Tag
Simon: コーディングエージェントを使う際、個人で使う方法なら理解できても、チーム環境でどう活用すればいいかについてはまだ明確にできていないという課題があります。Claude Tag は、まさにそのチームでの協働を担うレイヤーとしての回答ではないでしょうか。
Cat: その通りです。実際、セッションの多くがマルチプレイヤー形式で行われています。例えば「Cowork に新機能を追加すべきだ」と提案し、Claude Tag をタグ付けして第一ラウンドを担当させます。その後、「最終的な実装結果を記録として共有してほしい」と指示し、さらにデザイン担当者をタグ付けしてレビューしてもらいます。デザインチームからのフィードバックを受け取った後、エンジニアリングチームに引き継がれて完成まで進められ、本番環境へリリースされます。非常にスムーズな体験です。同じセッションをどう操るかという社会的なダイナミクスについてはまだ整理中ですが、人々は他者の使い方を観察し、そのノーム(慣習)に従う傾向があります。Claude Tag をチームに組み込む際も、直感的で問題なく受け入れられています。
Thariq: 人を教えるのにも役立ちますし、「スロップ(雑多な作業)」を減らす効果もあります。全員が一緒に Claude を使っている様子を目にするという事実自体が、Claude の使い方をレベルアップさせるからです。
これは、Midjourney が Discord チャンネルでパブリックにプロンプトを入力することを義務付けることで、高度な画像生成のプロンプト技術を普及させた事例を思い出させます。
ビルドコストが劇的に下がった今、どの機能に投資すべきかどう決めるのか?
私自身も非常に難しいと感じているのが、「コストが下がりすぎてしまった今、どの機能をリリースする価値があるのか」を見極めることです。
Simon: エンジニアリングにおける最大の難問である「優先順位の付け方」について、どうお考えですか?特に、機能の開発コストがこれほど安価になった現在、どの機能に投資し、リリースすべきかをどのように判断されているのでしょうか?
Cat: これがまさに難しいところです。私たちが取り組んでいるアプローチはいくつかあります。まず一つ目は、製品を毎日社内で実際に使い倒す「ドッグフーディング」です。社内製品で実現したい機能がない場合、代替手段を探すのではなく、その機能をサポートできるように製品自体を改善します。私たちは社内に非常に強いドッグフーディング文化があります。 世界中のユーザーに公開する前に、まずは Anthropic の全社員と、率直なフィードバック(できれば厳しい意見)を提供してくれる初期顧客に共有し、人々が本当に気に入るまで何度も改良を重ねます。製品を世に出すには、一定数のアクティブユーザー数やリテンション率という社内基準を満たす必要があります。 この基準が明確であるため、すべてのエンジニアが目指すべきゴールを理解しています。このプロセスは製品の完成度を高めることにもつながります。なぜなら、機能が洗練されていないとユーザーは離れてしまうからです。もしそうなれば、その機能はまだリリースするべきではないのです。
機能のリリース可否を、内部ユーザーの定着率という指標で判断するのは、私には非常に理にかなっていると思います。
驚いた機能の事例はありますか?
Simon: 驚いた機能の事例はありますか?リリースした当初、想定外のほどにエンゲージメントが爆発し、本来ならリリースされないだろうと見られていたものが、実際に製品として定着したようなケースです。
Cat: 確かに一つあります。チームの多くの方が リモートコントロール を愛用していますね。この機能を使えば、モバイル端末や Web ブラウザ上の Claude から、ローカル環境で CLI 上で動作している Claude Code セッションに接続できます。
私自身は、この機能が必要になることがありませんでした。というのも、私は簡単なコーディングタスクを直接モバイルで起動し、ローカル環境を使わずにクラウドセッションで実行するからです。そのため当初は「みんながリモート開発環境を設定すればいいのに」と思っていました。
しかし実際には、リモートコントロール機能をリリースした後に多くのユーザーにお話しすると、彼らが毎晩行っているのは、「ラップトップを充電器に繋ぎ、複数のリモートコントロールセッションを開き、画面をロックしてソファからスマホで Claude Code を操作する」という流れでした。これは私が当初理解していなかった、今ではチームが力を入れているワークフローなのです。
クロード・コードの生産環境コード、全行を人間がレビューしているのか?
今回のカンファレンスで繰り返し語られたテーマの一つに「レビュー」がありました。コーディング・エージェントが生成したコードを、人間がどの程度厳密にチェックしているかという点です。私はぜひとも、Claude Code の開発チームの考えを伺いたかったのです。
Simon: コードレビューはどのように行われているのでしょうか?生産環境にデプロイされる Claude Code 内の全行コードについて、人間によるレビューは必須となっていますか? もしそうではない場合、你们はどのような対策を講じて品質を維持しているのでしょうか?
Thariq: タスクの内容によって対応は異なります。重要な領域については「コードオーナー」を設けています。 システムプロンプトがその好例です。ここには明確なコードオーナーがおり、変更には必ず承認が必要です。
Simon: つまり、コードオーナーはその領域の品質に対して直接責任を負うわけですね。
Thariq: その通りです。
Cat: したがって、該当するコードに触れるプルリクエスト(PR)については、必ずそのオーナーの承認を得る必要があります。
Thariq: 私たちのチームでは、コードレビュー用の GitHub ボット がすべての PR をチェックしています。このボットはあらゆる PR に適用され、実際にはレビューの大部分を担っています。私がチームで目にするのは、より複雑な PR の場合、他のメンバーがレビューしやすいように「PR 解説用のアーティファクト」を作成するケースです。また、何かが失敗した際に必ずテストが実行されるよう、検証や CI/CD への投資にも力を入れています。Claude が Claude Code を制御してテストできる、非常に堅牢な環境も用意されています。つまり、コードレビューには多角的なアプローチを採用しているのです。
Cat: 一般的に、私たちは人間がループ(プロセス)に参加する必要のない世界へ移行しようとしています。Claude Code のコアや他のプロダクトのコアに対する最も重要な変更については、常にコードオーナーが存在し、すべての変更を手動でレビューします。しかし次第に、外側のレイヤーにおける変更については、Claude がコードレビューを完全に担当するようになりました。これは一見恐ろしく聞こえるかもしれませんが、ここに至るまでには 6 ヶ月以上ものプロセスがありました。コードレビューへの信頼を築くためには、小さな一歩(ベビーステップ)を踏むことが必要です。当初はすべての変更を手動でレビューしていましたが、次第に「このファイルに触れるコード変更については、コードレビューが 100% の問題を捕捉している」と判断し、「人間による手動レビューは不要だ」という方針へと移行していきました。
インシデント発生時のレビューでは、原因となった PR(プルリクエスト)を確認し、「どのようにしてコードレビューを改善すれば、同様の問題を検出できるか」を検討します。そしてその PR を評価セットに追加し、将来のコードレビューへの修正がその指標を低下させないことを保証しています。人間をコードレビューのループから外すことは大きな前進ですが、一晩でできることではありません。しかし、コードレビューが重要なすべての問題を捕捉しているという確信を得るために、インフラストラクチャに数ヶ月をかけて投資することで実現可能な道です。
つまり鍵となるのは、自動化プロセスを絶えず反復改善していくことです。
原文を表示
Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves.
The full video of the session is now available on YouTube. Below is an edited copy of the transcript, with extra links and my own bolded highlights.
*
A few top-level notes if you don't want to watch the video or wade through the whole transcript:
- Claude Tag (Claude's new collaborative Slack integration) now lands 65% of the product engineering PRs for the Claude Code team.
- Claude Code ships features to Anthropic employees first, and only ships the features that demonstrate user retention with that cohort
- Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the "outer layers" of the product.
- Adding examples to a system prompt is no longer best practice for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently reduced in size by 80%.
- Likewise, lists of "don't do X and don't do Y" can reduce the quality of results from the latest models.
- Dogfooding inside Anthropic is called "ant fooding".
- Anthropic really believe in their auto mode, and see that as an enabling technology for Claude Tag.
- Thariq advises offsetting coding-agent-induced Deep Blue by "being more ambitious" with the work you take on.
- Fable is competent at editing video, and Thariq used it to edit its own launch video.
- Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.
How has what you do day-to-day changed in the past year?
Simon: Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on the Claude Sonnet 3.7 launch. How has what you do on a day-to-day basis changed in the past year, now that we have these coding agents that actually work for us?
Cat: I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude. It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now.
Thariq: I remember the first text I got about Claude Code. One of my best friends was like, "You need to go try Claude Code." It was about when Opus 4 came out, and I tried it and I was like, "Oh, shit. I need to work at Anthropic now." And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that we have to do higher quality work than we've ever done before. The outputs are incredibly high quality. I've been using it to edit videos a bunch, and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before.
What piece of conventional software engineering no longer holds?
Simon: What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?
Cat: One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to develop more of your business sense and product sense on what it is we should build, because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's an increase in value on product taste and business sense, and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right.
Thariq: For me, it's that rewrites are now good.
Simon: The worst thing you could do is now actually fine!
Thariq: Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and I think the rewrite actually forces you to make sure you have a good test suite — but I think what people undercount is that a codebase is a spec, and maybe it's the only copy of the spec that you have, because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We rewrote Bun in Rust and it works great — it's live for me right now.
Simon: You're not shipping Claude Code on Bun-in-Rust yet, right?
Thariq: Internally we have.
(Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone on June 17th.)*
What kind of things are non-engineers doing with Claude Tag?
Simon: The other big launch recently was Claude Tag — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. What kind of things are non-engineers doing with Claude Tag?
Cat: Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about Claude Tag is it's multiplayer by default. Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, "Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase," and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that we've added team memory into this. If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team.
Internally, we see Claude Tag as the evolution of Claude Code. We see this as a large shift in how we work internally. Claude Tag currently lands 65% of our product eng PRs.
Simon: For all of Anthropic, or just for Claude Code?
Cat: This is just for our product engineering team — our internal version of Claude Tag lands 65% of our product PRs right now. And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. But Claude Tag is great for having it work proactively on your behalf, so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on.
Thariq: And for non-coding cases: for example, before this talk we asked Claude Tag, "Hey, when is Fable releasing?" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. As a search engine for your company, it's really valuable. It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, "Hey, tell me about this feature." They're not programmers, but Claude is a programmer — it can clone the codebase and say, "This is the feature, this is what it looks like, this is a recording of me using the feature." It enables a whole wide variety of things, and I think we're still early in figuring that out.
Claude Tag as the team collaborative layer
Simon: One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff.
Cat: Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, "Hey, I think we should implement this new feature in Cowork," and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, "Share a recording of your final implementation," and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session, but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams.
Thariq: It's great for teaching people, and also for reducing slop, because the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well.
This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels.
How do you decide which features are worth building when building is so much cheaper?
Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much.
Simon: How do you deal with the hardest problem in all of engineering — prioritization? How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now?
Cat: This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. We have a very heavy dogfooding culture internally. Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world. Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature.
Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me.
Do you have an example of a feature which surprised you?
Simon: Do you have an example of a feature which surprised you? You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing.
Cat: I do have one. A lot of folks on our team love remote control. Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, and then use their mobile phone from their couch to control Claude Code. So this has become a flow we're now leaning into that I didn't originally get — but now I do.
Does a human review every line of production code in Claude Code?
One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this!
Simon: How does code review work? Does a human being review every line of production code that makes it into Claude Code? And if not, what are you doing — how do you keep the quality up?
Thariq: It varies on the task a lot. For important areas we have code owners. The system prompt is an example where we have a code owner — you really need to get their approval.
Simon: So the code owner is directly responsible for the quality of that area of the code.
Thariq: That's right.
Cat: And they need to approve any PR that touches it.
Thariq: We have our code review GitHub bot review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that for more complex PRs you might make an artifact to explain the PR so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review.
Cat: In general, we are trying to move to a world where humans don't need to be in the loop. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, for the changes at the outer layers, we actually have Claude code review fully review those. That sounds pretty scary, but we've had a six-plus-month-long process to get here, and there are baby steps that you take to build up trust with code review. In the beginning we had human review for everything, and then increasingly we would say, okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those. And when we have incident review, we look at the PRs that caused the incident and say, okay, how do we update code review to catch that? — and we take those PRs and add them to an eval set to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do through many months of investment in the infrastructure to give you the confidence that code review is catching everything you care about.
So the key seems to be constantly iterating on the automate
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み