EU AI 法対応:AI 製品・エンジニアリングチームの実務ギャップ
本文の状態
日本語全文を表示中
詳細モードで約17分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Arize AI Blog
EU AI Act の施行に伴い、AI 製品チームは従来の原則方針を具体的な証拠へと転換する必要性に迫られ、評価と監査ログの整備が法的義務の実効性を決定づける。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 00:31
AI深層分析
キーポイント
原則から証拠への転換要求
EU AI Act は公平性や透明性といった抽象的な原則を、特定のシステムについて当局に提示可能な数値と履歴を持つ証拠へと変換することをチームに求めている。
評価プログラムの統合の必要性
RAI(責任ある AI)プログラムと評価プログラムは別個ではなく同一のものとして扱うべきであり、指標や所有者、閾値を両チームで合意する必要がある。
LLM 判定器の信頼性検証
バイアス判定に LLM 判定器を使用する場合、人間との合致率を検証した代表セットでの評価を行わなければ、誤った自信を生むリスクがある。
移行期間の誤解と注意点
市場に出ているシステムには移行期間が設けられるが、Article 111 の規定により、安易に待機する姿勢は推奨されず注意が必要である。
システム変更の記録とインストルメンテーションの重要性
AI エージェントは頻繁に変更されるため、事後に証拠を再構築するのではなく、作業中に履歴を自動記録する仕組みが必要である。
重要な引用
A principle written for your own organisation can stay a principle. A principle that a national authority might ask you to demonstrate, on a named system, on a Tuesday in 2028, has to turn into a number with a history behind it.
That's worse than having no metric, because it makes people confident when they shouldn't be.
This work is tedious, but it separates an operational programme from a decorative one.
Sixteen months of policy drafting will not create sixteen months of evidence. Instrumentation will.
編集コメントを表示
編集コメント
この記事は、規制対応を「コンプライアンス部門の課題」から「エンジニアリングチームの実践課題」へと位置づける視点を提供している。特に LLM 判定器の盲点に言及した点は、実務において即座に活用できる重要な洞察である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
欧州のチームは、同じ課題を繰り返し指摘しています。2〜3年前に策定された責任ある AI(RAI)ポリシーと、エージェントをリリースするエンジニアリングチームが存在するにもかかわらず、両者を結びつけるものがほとんどないというギャップです。EU AI 法はこのギャップを実務的な問題へと変えます。チームは、特定のシステムが時間とともにどのように振る舞い、変化し、レビューされたかを示す必要があるかもしれません。
開発者にとっては、トレーサビリティ(追跡可能性)と評価が求められます。プロダクトマネージャーにとっては、責任者の明確化、閾値の設定、そして証拠に基づいたリリース判断が重要になります。
免責事項:これはエンジニアリングの視点から私が行った独自の分析であり、法的助言や公式ガイドライン、あるいは Arize 社の公式見解を表すものではありません。以下に記載する日付と義務は、2026 年 8 月時点における EU AI 法およびデジタル・オムニバス条令に対する私の解釈に基づいています。詳細は依然として流動的であるため、行動を起こす前に必ず『欧州連合公報』や自社の法務担当者と照合してください。
実務的な変化:原則から証拠へ
第 III 章を読み進めても、これまで公表したことがないような新たな価値観を見つけることはできません。公平性、透明性、人間の監督、堅牢性、プライバシー、説明責任。おそらくあなたのポリシーにもこれらの用語が 5 つほど含まれているはずです。
異なる点は、その対象となる読者が誰かです。組織内部向けに書かれた原則は、単なる理念として留まることができます。しかし、2028 年の火曜日に国家当局から特定のシステムについて実証を求められるような原則であれば、背後にある履歴を持つ数値へと具体化されなければなりません。
この 6 つの原則を実際の運用アーティファクトに落とし込むための一つのアプローチがあります。その根拠となる証拠の多くは、AI エンジニアリングの日常から得られるものです。具体的には、トレース(追跡記録)、評価結果、注釈付きデータ、CI(継続的インテグレーション)システム、監査ログなどが該当します。

図の右側の列に示されているのは、特別なコンプライアンス専用ツールではありません。むしろ、評価を真面目に取り組むチームが日常的に使用する標準的な設備です。だからこそ私は、リスク管理とAI倫理(RAI)プログラムと評価プログラムは一体となるべきだと主張し続けています。製品・エンジニアリングチームは、各原則に対して「どの指標を使うか」「責任者は誰か」「閾値はいくらか」「レビューのフローはどうするか」「リリースにどう影響するか」について合意しておく必要があります。
指標を信頼する前に知っておくべき注意点
何度も目にしてきた失敗パターンがあるので、あえて率直にお伝えします。ある人が LLM(大規模言語モデル)による判定器を立てて、バイアス基準に対してスコアリングを行い、その結果をリスク委員会の資料に載せるケースです。しかし、実際に重要なケースにおいて、その判定器が人間の判断と一致しているかどうかを検証した人は誰もいません。
これは指標がない状態よりも悪いです。なぜなら、本来は自信を持つべきではない状況で、人々が根拠のない自信を持ってしまうからです。公平性スコアをチームから出す前に、代表的なキャリブレーションセットを用いて「判定器と人間の一致度」を測定し、不一致となったケースを精査してください。さらに、モデルやトラフィックがドリフト(変化)した際にも参照できるよう、そのセットは保存しておく必要があります。単純な精度だけでは不十分です。この作業は地味で手間がかかりますが、それが運用可能なプログラムと単なる飾り付けのプログラムを分ける決定的な違いとなります。
なぜあと 16 ヶ月待てばいいという理由にはならないのか
オムニバス法(EU AI 法)の表面的な読み方では、Annex III に該当するチームは 2027 年後半まで猶予があるように思える。しかし私を不安にさせるのは第 111 条だ。
すでに市場に出ているシステムには移行期間が設けられるが、これは「大幅な改変」が行われていない場合に限られる。この規定は、年に数回しかリリースしないソフトウェアを想定して作られたものだ。AI エージェントについては記述されていない。

チームに問うべきだ。「直近の四半期で、システムプロンプトを変更し、検索インデックスを再構築し、より新しいモデルへ移行し、ツールを 2 つ追加した。これらの中でどれが『大幅な改変』にあたるか?」と。多くのチームは答えられない。なぜなら、誰が見ても確認できる形式で変更履歴が記録されていないからだ。
私がこれまで見てきた一貫したアドバイスは、「猶予期間を待つのではなく、その時間を活用せよ」というものだ。16 ヶ月間の政策策定プロセスが、自動的に 16 ヶ月分の証拠資料を生み出すわけではない。証拠となるのは、システムに組み込まれた計測機能(インストゥメンテーション)だけだ。
ドキュメントをシステムの一部にする
リスクが高いとされる章の条文は、エンジニアとして読み直せば法律用語ではなく技術要件に見える。第 12 条は記録の保持を求め、第 14 条はシステムを理解し介入できる人材の配置を要求する。第 15 条は、検証可能な精度と堅牢性を求める。そして第 72 条は、リリース後も継続的な監視を義務付ける。
これらはすべて、テレメトリ(計測)、評価、ワークフロー、リリース管理に関する要件だ。一度きりの事後対応で再現するのは不可能だ。なぜなら、この法は「履歴」そのものを求めているからだ。システムが作業が行われている最中に、その履歴を書き残す仕組みが必要となる。
義務とそれに応えるアーティファクト
Obligation
The artifact that answers it
第 12 条(記録の保存)
エージェントの全軌跡における OpenInference/OpenTelemetry のトレースを記録します。モデル呼び出し、ツール引数、検索結果、ハンドオフなどです。保持期間は意図的に選定したものです。
第 14 条(人間の監督)
共有スキーマに基づくスパンへの注釈と、失敗をレビュー担当者にルーティングするキューを設置します。各ラベルには、担当者名と時刻を含めます。
第 15 条(精度と堅牢性)
評価はオフラインのデータセット上で行う一方、オンラインでは実稼働トラフィックに対して実施し、CI に組み込んで回帰問題がユーザーに届かないようにします。
第 10 条(データガバナンス・バイアス検証)
外部へエクスポートする前に、自社のプロセス内でマスキングと PII の削除を行います。評価は学習セットだけでなく、実稼働トラフィック上の不均衡指標も対象とします。
第 50 条(透明性)
開示が実際に存在したことを確認する評価者を用意します。ポリシーで開示を義務付け、その実施を示す仕組みです。
第 72 条(市場投入後の監視)
リリースのゲートに用いる指標と同一のメトリクスに対して監視とアラートを設定し、デプロイ前後で品質基準が統一されるようにします。
第 11 条および附属書 IV(技術文書)
バージョン管理されたプロンプト、データセットの系譜、実験結果を記録します。技術ファイルの変更履歴は、要約として機能するものです。
一般的な注意:これはエンジニアリング向けのガイダンスです。システムがどのように分類され、どの適合性ルートを選択すべきかは、法務担当者と協議すべき事項です。
証拠のための参照アーキテクチャ
上記のマッピングが機能するのは、基盤となる仕組みが信頼性が高く検証可能な証拠を生成する場合に限られます。Arize AX を用いた場合の参照アーキテクチャは以下の通りです。

エージェントのランタイム、アプリケーション内で動作するデータ削除(redacting)スパンプロセッサ、OTel カスタマー、ストレージ(EU 地域または自社のクラスター)、オンラインおよびオフラインの評価器、失敗事例のラベリングキュー、ベンチマークデータセット、次のリリース前の CI ゲート、そしてモニタリングと監査ログ。
このフローにおいて、3 つ目のコンポーネントの位置づけを誤解している人が少なくありません。もしデータ削除を遠隔側ではなく、自社のプロセス内で実施すれば、「完全な記録」と「データの最小化」の間で二者択一を迫られる必要はなくなります。両立が可能になるのです。
- レイヤー:動作内容 / 残されるもの / 原則
- 1. インストルメンテーション:OpenInference と OpenTelemetry のインストルメンテーションに加え、スパンプロセッサが属性クラスをマスクし、正規表現または Presidio を用いて個人情報を削除する / すべての実行記録は完全であると同時に最小化されている / プライバシー(第 12 条、第 10 条)
- 2. ストレージと居住性:Arize AX の EU リージョン(ベルギー)、または自社クラスター上の自己ホスト型 AX とオブジェクトストレージ、あるいはエアギャップされたインストール / 特定の管轄権下にある履歴データ。監査の境界は一つに限定される / プライバシー(第 12 条、GDPR)
- 3. 評価:生トラフィックに対するオンライン評価器、データセットを対象としたオフライン評価器、決定論的チェックと LLM によるジャッジを併用。ジャッジと人間の合意率も測定する / 時間経過に伴う公平性、安全性、根拠の明確さ、開示に関する指標。その信頼性を裏付ける根拠がある / 公平性、堅牢性、透明性(第 15 条、第 10 条、第 50 条)
- 人間のレビュー
スキーマを定義するアノテーション設定、失敗時に特定のレビュアーへルーティングするキュー、誰が何をレビューし、どのような判断を下し、いつ行ったのかという記録。人間による監督(14)
- 変更管理
CI(GitHub Actions、GitLab、Jenkins、Azure DevOps)におけるベンチマークデータセット、プロダクションモニタリング、アクセスとエクスポートの監査ログ。リリース履歴には理由、ドリフトアラート、アクセスの追跡記録を含める必要があります。説明責任(15, 72, 11)
ここでは、意図的に設定すべき2つのデフォルトがあります。
保持期間:デバッグ用のデフォルトは短期間の記憶を前提としています。しかし、規制上の証拠保全には明確な期間の定義が必要です。ケースに適用される義務に合わせて、トレースとアノテーションの保持期間を設定し、直近のクォーターのトレースに頼る前に、アノテーションの遡及範囲が適切か確認してください。
サンプリング:観測可能性(オバザビリティ)においてはサンプリングは一般的ですが、記録管理の観点からは危険を伴います。リスクの低いトラフィックについては適切な場合にサンプリングを行ってください。ただし、高リスクのワークフローは、リスク担当者と法務担当者が別のアプローチを承認しない限り、完全な忠実度(フルファイドリティ)で保持してください。5% のトラフィックに基づいて計算された公平性スコアが有用である場合もありますが、それは完全な記録とは異なります。
エージェント1つでエンドツーエンドを実現する
信用力評価アシスタントは良い例です。これは明確に付録III に該当し、かつ責任あるAIにおける最も古くからの課題の一つでもあります。つまり、誰かの資金アクセスに関する自動化された決定を、モデルによって説明するという問題です。
顧客が「なぜ自分の利用限度額がこのように設定されたのか」と尋ねたと仮定しましょう。エージェントはポリシー文書を取得し、スコアリングサービスに問い合わせ、回答の草案を作成します。周囲のワークフローには以下の5 つの機能が必要です:
トレースは、名前や口座番号が置換された状態で既に匿名化されて記録されます。プロンプト、取得したポリシーテキスト、ツール呼び出し、出力などすべてが、スパンがプロセスから離れる前に処理されるため、転送中の個人データは一切含まれません。
評価者は、実際に取得されたポリシーに対する整合性、申請者グループ間の結果の格差、Article 50 の開示が存在したかどうかをスコアリングします。
不具合は人間に引き継がれます。整合性が低いトレースはキューに送られ、アナリストが説明を確認または覆し、その際に自分の名前とスパンへのタイムスタンプが付与されます。これが Article 14 の証拠であり、注意すべき点は、誰かが証拠として意図的に生成したわけではないということです。彼らは単に職務を遂行する過程で自然にそれを生み出したのです。
不具合は次のステップのゲートとなります。レビュー済みのトレースはベンチマークデータセットとなり、次回のプロンプトやモデル変更はこのデータセットに対して CI で実行されます。回帰テストが失敗すればビルドも失敗し、つまりすべてのリリースには明確な理由が文書化されることになります。
モニタリングは継続されます。同じ指標におけるドリフトを検知してアラートが発令され、ログイン、変更、データのエクスポート(どのプロジェクトから誰が取得したかを含む)を網羅する監査ログが維持されます。

規制当局のためだけに存在するステップはありません。これは、エージェントの信頼性を高め、製品の運用を容易にするために構築する同じループそのものです。
何かを計測(インストゥルメント)する前に、データがどこに保存されるかを決定してください。
信用取引に関する会話のトレースは個人データです。AI Act への対応のために GDPR の問題を生み出すような解決策では、進歩とは言えません。
Arize AX は、制御の度合いが順に高まる 3 つのデプロイメントパターンをサポートしています。1 つ目はベルギーにある EU リージョンです。2 つ目は、Kubernetes クラスターとオブジェクトストレージ上でセルフホスティングする方式で、この場合 Arize は一切データを保持しません。3 つ目は、機密情報や厳格な規制が適用される環境向けに、外部への通信経路を遮断したエアギャップ型インストールです。どのパターンを選ぶかは、データの所在(レジデンシー)、運用上の責任の所在、そして監査すべきセキュリティ境界に影響します。
これら 3 つのパターンすべてにおいて、SOC 2 Type II、ISO 27001、PCI DSS、HIPAA、GDPR の認証を取得済みです。また、SAML 2.0 SSO をご自身の IdP と連携させたり、プロジェクト単位まで細かく設定できるロールベースアクセス制御も提供しています。

Arize AX が行わないこと
これらの責任の一部は、他者が担うべきものです。当社は、お客様のシステムが「ハイリスク」に該当するかどうかを決定する立場ではありませんし、適合性評価の実施や CE マークの付与、登録手続きの代行、あるいは第 17 条で求められる品質マネジメントシステムの構築も行っていません。これらは誰かが判断を下さなければならない事項であり、数値化できるものでもありません。これらの責任は、貴社のガバナンス担当者や法務チームが負うべきものです。
基本権評価もその対象に含まれますが、ここでは少し補足させてください。私が目にしてきた第 27 条に基づく評価の多くは、「システムが何をするだろう」という人々の推測に基づいて作成されています。しかし、実際のデータは「差異の数値」や「レビュー済みのトレース」です。最終的な判断を下すのは貴社ですが、何を判断しているのかを推測する必要はありません。
まだ実現できていないこともあります。改ざん防止が可能な記録を提供することはできません。トレーシング、評価結果、注釈は監査担当者に「何が起きたか」「誰が承認したか」を示すことはできますが、後からファイルが編集されていないことを証明する手段はまだ確立されていません。そのための基準も現在策定中です。監査の最中に気づいて慌てるより、今からお伝えしておきます。
そして、どのベンダーも提供できない領域があります。ガバナンスプラットフォームはリスク登録簿やアテステーションを整理整頓してくれますが、「火曜日の午前4時12分にエージェントが何をしたか」「誰がそれを確認したか」「先週のプロンプト変更が悪影響をもたらしたかどうか」まで教えてくれるわけではありません。そのような詳細な事象を把握できるのは、自社の計測システムだけです。
「雰囲気」だけで製品を出荷しないでください。特に規制当局に対してはなおさらです。
2027 年 12 月までは、トレーシングデータを保持しているかどうかが問われる 16 ヶ月間となります。まずは Annex III に該当するユースケースの一つから始めましょう。今日でもっとも証明が難しいとされる原則を特定し、評価指標と責任者を定義した上でワークフローに計測機能を組み込み、その結果をリリースプロセスに組み込んでください。
デモ予約 | Self-host Arize AX | SaaS 版 Arize AX | トレーシングドキュメントの閲覧
これはエンジニアリングの実践に関する情報であり、法的助言ではありません。本記事に記載の日付は、2026 年 7 月 27 日現在の AI デジタルオムニバス法に基づくものです。信頼する前に、官報や自社の弁護士と照合することをお勧めします。
この投稿「AI プロダクトおよびエンジニアリングチームのための EU AI 法の解明」は、Arize AI で最初に公開されました。
原文を表示
European teams keep describing the same gap: a Responsible AI (RAI) policy written two or three years ago, an engineering team shipping agents, and almost nothing connecting the two. The EU AI Act makes that gap operational: teams may need to show how a named system behaved, changed, and was reviewed over time. For developers, that means traces and evaluations. For product managers, it means owners, thresholds, and release decisions backed by evidence.
Disclaimer: This is my own analysis, written from an engineering perspective. It is not legal advice, formal guidance, or a statement of Arize’s position. The dates and obligations below reflect my reading of the EU AI Act and the Digital Omnibus as of August 2026, and the details are still moving. Verify anything you plan to act on against the Official Journal and with your own counsel.
The practical shift: principles become evidence
Read Chapter III and you won’t find any values you haven’t already published. Fairness, transparency, human oversight, robustness, privacy, accountability. Your policy probably uses five of those words.
What’s different is who the audience is. A principle written for your own organisation can stay a principle. A principle that a national authority might ask you to demonstrate, on a named system, on a Tuesday in 2028, has to turn into a number with a history behind it.
Here is one way to translate those six principles into operational artifacts. Most of the evidence comes from ordinary AI engineering: traces, evaluations, annotations, CI, and audit logs.

Most of the right-hand column is not specialized compliance tooling. It is the ordinary equipment of a team that takes evaluations seriously. That’s why I keep arguing that the RAI programme and the evaluation programme should be one program. Product and engineering teams should agree on the metric, owner, threshold, review path, and release consequence for each principle.
One warning before trusting the metrics
There’s a failure mode I’ve now seen enough times to be blunt about it. Someone stands up an LLM judge, points it at a bias criterion, gets a score, and puts that score in a risk committee deck. Nobody ever checked whether the judge agrees with a human on the cases that actually matter.
That’s worse than having no metric, because it makes people confident when they shouldn’t be. Before a fairness score leaves your team, measure judge-to-human agreement on a representative calibration set, inspect the disagreement cases, and retain that set as models and traffic drift. Raw accuracy alone is not enough. This work is tedious, but it separates an operational programme from a decorative one.
Why sixteen more months are not a reason to wait
The obvious reading of the Omnibus is that Annex III teams can relax until late 2027. The bit that worries me is Article 111.
Systems already on the market get a transition period, but only if they aren’t significantly modified. That was drafted with the assumption of software that ships a couple of times a year. It doesn’t describe an agent.

So ask your team this: over the last quarter you changed the system prompt, rebuilt the retrieval index, moved to a newer model, and added two tools. Which of those was the significant modification? Most teams can’t answer because nobody recorded the changes in a form anyone can review.
The consistent advice I have seen is to use the extra time, not wait it out. Sixteen months of policy drafting will not create sixteen months of evidence. Instrumentation will.
Make documentation part of the system
Most of the high-risk chapter stops sounding legal once you read it as an engineer. Article 12 wants records. Article 14 wants a person who can understand the system and step in. Article 15 wants accuracy and robustness you can show. Article 72 wants you to keep looking after release.
These are telemetry, evaluation, workflow, and release-management requirements. You cannot reconstruct them reliably after the fact because the Act asks for a history. The system has to write that history while work happens.
Obligation
The artifact that answers it
Art. 12 Record-keeping
OpenInference/OpenTelemetry traces across the whole agent trajectory: model calls, tool arguments, retrieval results, handoffs. With retention you’ve chosen deliberately.
Art. 14 Human oversight
Annotations on spans using a shared schema, plus queues that route failures to reviewers. Each label carries the reviewer and the time.
Art. 15 Accuracy and robustness
Evaluators offline on datasets and online against live traffic, wired into CI so a regression doesn’t reach users.
Art. 10 Data governance, bias examination
Masking and PII redaction inside your own process, before anything is exported. Disparity metrics on production traffic, not only on the training set.
Art. 50 Transparency
An evaluator confirming the disclosure was actually there. A policy requires disclosure; this shows it happened.
Art. 72 Post-market monitoring
Monitors and alerts on the same metrics you gate releases with, so quality means one thing before and after deployment.
Art. 11 / Annex IV Technical documentation
Versioned prompts, dataset lineage, experiment results. The change history of your technical file is meant to summarise.
Usual caveat: this is engineering guidance. How your systems get classified, and which conformity route you take, is a conversation for your counsel.
A reference architecture for evidence
The mapping above only works if the plumbing produces reliable, reviewable evidence. With Arize AX, the reference architecture looks like this.

Agent runtime, then a redacting span processor running inside your application, then an OTel collector, then storage (EU region or your own cluster), then evaluators online and offline, then a labeling queue for the failures, then a benchmark dataset, then a CI gate on the next release, then monitors and audit logs.
The position of that third component is the part people get wrong. If redaction happens in your process rather than at the far end, you stop having to choose between a complete record and data minimisation. You get both.
Layer
What runs there
What it leaves behind
Principle
- Instrumentation
OpenInference and OpenTelemetry instrumentation, plus a span processor that masks attribute classes and redacts PII by regex or with Presidio
A complete but minimised record of every run
Privacy (12, 10)
- Storage and residency
The Arize AX EU region in Belgium, or self-hosted AX on your own cluster and object storage, or an air-gapped install
History under a known jurisdiction, one boundary to audit
Privacy (12, GDPR)
- Evaluation
Online evaluators on live traffic, offline evaluators over datasets, deterministic checks alongside LLM judges, with judge-to-human agreement measured
Fairness, safety, groundedness and disclosure metrics over time, and a reason to believe them
Fairness, robustness, transparency (15, 10, 50)
- Human review
Annotation configs defining the schema, queues routing failures to named reviewers
Who reviewed what, what they decided, when
Human oversight (14)
- Change control
Benchmark datasets in CI (GitHub Actions, GitLab, Jenkins, Azure DevOps), production monitors, audit logs on access and export
A release history with reasons, drift alerts, an access trail
Accountability (15, 72, 11)
Here are two defaults to set deliberately:
Retention. Debugging defaults assume a short memory. Regulatory evidence requires a deliberate horizon. Set trace and annotation retention against the obligations that apply to the use case, and verify the annotation lookback window before relying on last quarter’s traces.
Sampling. Sampling is normal in observability but dangerous for record-keeping. Sample low-risk traffic where appropriate. Keep high-risk workflows at full fidelity unless your risk and legal owners approve another approach. A fairness score computed on 5% of traffic may be useful, but it is not the same as a complete record.
One agent, end to end
A creditworthiness assistant is a good example because it’s squarely Annex III and it’s also the oldest problem in Responsible AI: an automated decision about someone’s access to money, explained by a model.
Suppose a customer asks why their limit was set where it was. The agent retrieves policy documents, calls a scoring service, and drafts an explanation. The surrounding workflow should do five things:
The trace is captured, already redacted. Prompt, retrieved policy text, tool call, output, all of it, with names and account numbers replaced before the span leaves the process. Complete record, no personal data in transit.
Evaluators score it. Groundedness against the policy that was actually retrieved. Outcome disparity across applicant cohorts. Whether the Article 50 disclosure was present.
Failures go to a person. Low-groundedness traces land in a queue. An analyst confirms or overturns the explanation, and their name and the timestamp attached to the span. That’s the Article 14 evidence, and note that nobody generated it as evidence. They generated it by doing their job.
The failures become the gate. Reviewed traces turn into a benchmark dataset, and the next prompt or model change runs against it in CI. A regression fails the build, which means every release has a documented reason behind it.
Monitoring continues. Alerts on drift in the same metrics. Audit logs covering logins, changes, and every export, including which project the data came from and who pulled it.

No step exists only for the regulator. It is the same loop you would build to make the agent more reliable and the product easier to operate.
Decide where the data lives before you instrument anything
Traces of a credit conversation are personal data. So, solving your AI Act problem by creating a GDPR problem isn’t progress.
Arize AX supports three deployment patterns, in increasing order of control: an EU region in Belgium; self-hosting on your Kubernetes cluster and object storage, where Arize stores nothing; and an air-gapped install with no outbound path for classified or heavily regulated environments. The choice affects residency, operational ownership, and the security boundary you must audit.
Across all three: SOC 2 Type II, ISO 27001, PCI DSS, HIPAA, GDPR, SAML 2.0 SSO against your own IdP, role-based access down to individual projects.

What Arize AX doesn’t do
Some of this belongs to other people. We don’t decide whether your system counts as high-risk. We don’t run your conformity assessment, put a CE mark on anything, file your registration, or build the quality management system Article 17 asks for. Those are calls someone has to make, not things you can measure, and they sit with your governance people and your lawyers.
Your fundamental rights assessment is on that list too, though I want to push on it a bit. Most of the Article 27 assessments I’ve seen are built on what people assume the system does. The disparity numbers and the reviewed traces are what it actually did. You still have to make the judgement call. You just don’t have to guess at what you’re judging.
There’s also something we genuinely can’t do yet. We don’t give you tamper-proof records. Traces, evals and annotations will show an auditor what happened and who signed off, but nobody can prove afterwards that the file wasn’t edited. The standards for that are still being written. I’d rather tell you now than have you work it out halfway through an audit.
And then there’s the bit no vendor can sell you at all. A governance platform will keep your risk register and your attestations tidy. It won’t tell you what your agent did at 04:12 on a Tuesday, whether anyone looked at it, or whether last week’s prompt change made things worse. Only your own instrumentation knows that.
Don’t ship vibes. Least of all to a regulator.
December 2027 is sixteen months of traces you will either have or you will not. Start with one Annex III use case. Identify the principle you would struggle most to demonstrate today, define the metric and owner, instrument the workflow, and make the result part of the release process.
Book a demo · Self-host Arize AX · Arize AX on SaaS · Read the tracing docs
Engineering practice, not legal advice. The dates here reflect the Digital Omnibus on AI as in force on 27 July 2026 and are worth checking against the Official Journal and your own counsel before you rely on them.
The post Demystifying the EU AI Act for AI product and engineering teams appeared first on Arize AI.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み