Hugging Face、ICML 論文 2,200 件の再現ハッカソンで得た教訓を報告
本文の状態
日本語全文を表示中
詳細モードで約14分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Hugging Face Blog
Hugging Face は、コミュニティの1,200名以上がコードエージェントを用いてICML 2026の論文2,226篇を再現したハッカソンの結果を発表し、AI による研究実験における人間の役割について考察している。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月14日 02:54
AI深層分析
キーポイント
大規模な自動再現プロジェクトの実施
19日間のハッカソンにおいて、コミュニティの参加者が自身のコードエージェントを用いて2,226篇の論文を再現し、6,816件のトラックログが提出された。
AI エージェントによる研究検証の可能性
この実験は、AI エージェントが個別の主張ごとに論文を再現・検証する能力を実証し、大規模な科学論文の再検証プロセスへの応用を示唆している。
人間と AI の協働における役割の変化
記事は、AI エージェントが実験作業を担う未来において、研究者やコミュニティが果たすべき新しい役割について考察を行っている。
AI 研究の提出数とレビュー能力の乖離
ICML 2026 の採択数は前年の約倍であるが、ボランティアであるレビュアーのレビュー能力は追いついていない。
コードエージェントによる大規模再現の可能性
Claude Code や Cursor などのコーディングエージェントを用いれば、人間が週末を要する検証を並列処理で半日以内に行える。
重要な引用
In this post, we're sharing what we learned from running this hackathon, and what it suggests about the role humans will play when agents are doing the research experiments.
participants published 6,816 Trackio logbooks reproducing 2,226 papers, about a third of the conference
"My low confidence score is because I did not check all the proofs carefully."
if we actually re-examined a major conference at scale, and tried to reproduce every paper, what would we find?
編集コメントを表示
編集コメント
ICML 2026という未来の時点でのイベントを扱っている点は興味深いが、これはハッカソンの設定やシナリオとして提示されていると解釈される。AI エージェントが学術研究の検証プロセスに参画する具体的な事例を示しており、今後の研究手法の変容を考える上で示唆に富む内容である。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
「記事一覧に戻る」
- 誰もが多角的にレビューしきれない論文の数々
- ハッカソンの開催(7 月 15 日〜8 月 2 日)
- 発見された事実と検証結果 高品質な再現事例
- 誤りや偽装の発覚、そしてその後の対応
- 著者との対話
- 人間に求められる役割
7 月、私たちはコミュニティメンバー 1,200 名以上が参加するハッカソンを主催しました。彼らは各自のコーディングエージェントを持ち込み、ICML 2026 で発表された論文を主張ごとに再現しようと挑戦しました。わずか 19 日間で、参加者は 2,226 編の論文(学会全体の約 3 割)について 6,816 件の Trackio ログブックを投稿し、その成果を公開しました。
本稿では、このハッカソンを通じて得られた知見と、エージェントが研究実験を担う未来において人間が果たすべき役割について考察します。
誰もが多角的にレビューしきれない論文の数々
AI 研究の再現性に関する疑問は、現在の AI ブーム以前から存在していました。しかし、規模が拡大するにつれてこれらの疑問はより深刻化しています。
ICML 2026 では 23,918 件の論文が提出され、そのうち 6,352 件が採択されました。これは前年の約倍に相当し、AI エージェントによる実験実行や論文執筆の高速化が要因の一つとなって、指数関数的な増加傾向が続いています。
一方、査読のキャパシティはそれに追いついていません。主要カンファレンスの査読者はボランティアであり、論文を完全に審査するための時間や専門知識を持っていないケースも少なくありません。ある ICML 2026 のスポットライト採択論文に対する査読者の言葉をご紹介します。
「私の低い評価点は、証明部分をすべて慎重に確認しなかったためです。」
この論文は高い評価を受け、スポットライトにも選ばれました。 これは後ほど再度取り上げる重要なポイントなので、覚えておいてください。実際に証明を慎重に検証した際に何が起きたかについても触れます。
しかし、状況が変わったのは、提出数の増加を招いた技術が、その増加に対処する手助けもしてくれるようになった点です。Claude Code、Codex、Cursor、Pi といったコーディングエージェントを使えば、論文を読み込み、コードを書き、実験を実行し、結果を報告することまで自動化できます。以前は査読者が週末を費やして行う必要があった詳細な確認作業も、エージェントなら午後数時間で並列に数千回試すことが可能です。
そこで私たちが問いかけたかったのはこうです。主要カンファレンスを大規模に再検討し、すべての論文の再現を試みた場合、何が明らかになるのか?
ハッカソン(7 月 15 日〜8 月 2 日)
私たちが直接論文を検証するのではなく、コミュニティ全体に開放しました。これにより、エージェント・フレームワークの多様性、計算リソースの制約、そして科学への感性まで、あらゆる要素が参加者から集まります。
2026 年 7 月 15 日から 8 月 2 日にかけて行われた「ICML 2026 オープン再現チャレンジ」[https://huggingface.co/spaces/ICML-2026-agent-repro/challenge] は、以下のような仕組みで運営されました。
- 論文を選ぶ。ICML 2026 に採択された 6,341 編の論文すべてをインデックス化し、各論文の要約と核心的な科学的主張を抽出しました。これにより、エージェントは 40 ページに及ぶ PDF を読むのではなく、具体的で検証可能な目標から作業を開始できます。同じ論文を複数の参加者が再現することも推奨されました。
- 自分自身のアージェントを用意する。参加者は Claude Code、Codex、Cursor、OpenResearch の
orxなど、さまざまなツールやその中間のものを活用しました。エージェントが単一のコマンドで論文、主張、そしてチャレンジの指示を取得できるよう、簡素化されたインターフェースを提供しました。
- 再現し、すべてを公開する。すべての試行結果は Trackio のログブックとして記録されます。これは、解説文、実行コード、生成された成果物を含む静的な Hugging Face Space です(オプションで、エージェントの実行履歴全体も Hugging Face Dataset としてアップロードされます)。監査プロセス自体も、監査可能である必要があります。
審査を行う。自動化されたログブック審査員(オープンウェイトモデルの GLM-5.2 を実行)がすべてのログブックを再読し、各主張に対して「検証済み」「反証済み」「玩具級(証拠は縮小規模)」、「結論未定」の判定を下しました。この審査員には、各ログブック内の自己評価を信頼しないよう明確に指示されていました。
参加者には Hugging Face の計算リソースとして 20 ドル分のクレジットが付与され、HF Jobs で実験を行いました。今回のチャレンジ全体で 2,962 件のクラウドジョブが起動されました。論文のデータセットが専用であったり、チェックポイントが公開されていなかったりして完全な再現が不可能な場合、参加者は元の論文の特性を模倣した合成データ上で玩具級の実験を行いました。
完成した再現事例は以下のようになります。
数値で見ると、このハッカソンはおそらく科学会議における最大の試みとなる再現プロジェクトでした:
- 1,221 人のコミュニティメンバーが 組織 に参加しました。
- 6,816 件の再現ログブックが公開されました。
- 2,226 篇の論文に挑戦し、これは会議全体の 34% を占めます。多くの論文には複数の独立したチームが取り組んでいました。
- 35,908 件もの主張が審査され、すべての判定はチャレンジ終了時に公開データセットとして凍結されました。
- 2,962 件の HF Jobs が起動され、そのうち 274 件の完全なエージェント・トレースデータセットが Hugging Face に公開されています。
私たちが発見したこと
論文ごとの主張レベルの判定を集計すると:
調査対象となった論文の51%(1,103本)には、少なくとも一つの実験で独立して検証可能な主張が含まれていました。そのうち266本は完全再現され、抽出されたすべての主張が確認されました。さらに632本は部分的に再現され、偽証された主張は見つかりませんでした。合計で3,978件の個別の主張が、実際の実験によって裏付けられました。
一方、調査対象となった論文の23%(496本)には、少なくとも一つの実験で否定または争われるべき主張が含まれていました。これには、すべての主張が偽証され、検証可能なものが一つもない49本の論文も含まれます。最も興味深いのは、独立した再現チームが同じ主張に対して*逆の結論*を下した242本もの論文があることです。再現可能性は二値(Yes/No)ではなく、対立的なプロセスなのです。
残りの論文はその中間に位置しました。502本は玩具規模の実験データしかなく、280本は何らかの結論を出すことができませんでした(最も一般的な原因は、必要なアーティファクトが欠落していることです)。
再現が成功した事例
いくつかの論文はこの試練をくぐり抜け、見事な結果を残しました。コミュニティが蓄積した最良の記録は、それ自体として読む価値があります:
- 「Flat Minima and Generalization: Insights from Stochastic Convex Optimization」 は、20の独立したチームによって再現されました。そのうち12チームがすべての主張を検証しました。リンク先の記録には、エージェントの完全な実行トレースが含まれており、公開されています。
「A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness」(https://huggingface.co/spaces/gchauhan/repro-a-coin-flip-for-safety-llm-judges-fail-to-reliably-measure-adversarial-robustness) について、17 の検証記録のうち 14 がすべての主張を確認しました。信頼できない LLM 判定器を扱った論文が、LLM エージェントによる厳密な検証にも耐えうるのかどうかという興味深いケースです。
捏造の発見と、その後の展開
35 人の参加者が何らかの捏造を公式に主張しました。私たちは、主張されたすべての捏造事例に対して敵対的な再検証を行いました。具体的には、論文と検証記録(ログブック)を再読し、論文本文から数式の導出や実験の実装をやり直したのです。
以下は、確認された捏造事例の一部で、それぞれ発見した検証記録へのリンクを記載します。
序論に登場するページングに関する論文。 証明書を慎重にチェックしなかった審査員の問題でしょうか。論文「Towards Optimal Robustness in Learning-Augmented Paging」では、提案されたアルゴリズムが Hk+O(1) の堅牢性を達成すると主張されています。しかし、ある参加者の検証記録 (https://huggingface.co/spaces/Auenchanters/repro-towards-optimal-robustness-in-learning-augmented-paging) では、付加項が 0.38lnk のように成長していることが測定され、証明が破綻する具体的なステップが特定されました。私たちの再実装では k = 1,024 まで範囲を広げ、約 9 シグマの成長を確認しました。真の堅牢性は Hk+Θ(logk) です。
224 ステップ後に崩壊する定理。「Attention の順伝播と Frank-Wolfe 法」は、原点がトークン粒子の凸包内に含まれる場合、それらが原点に収束することを証明しています。しかし、3 つの独立したチームがこの主張に対する反例を見つけました。違反はそれぞれ t = 224、約 3,800、6,416 ステップで初めて確認されています。これが、なぜ他の多くの検証者がこの定理を「正しい」と判断してしまったのかを説明しています。彼らは有限の時間範囲内でのチェックを行っただけだったため、問題に気づく前に検証が終了してしまっていたのです。
最も明確な反例 The cleanest counterexample は有理数算術で厳密に記述されているため、浮動小数点の誤差を隠れ蓑にする余地はありません。著者たちは同日にこの反例を確認し、修正に取り組んでいます。
「一つの損失関数向けに書かれた理論が、別の損失関数によって生み出された結果」という皮肉な現象も報告されています。「Self-Distillation Enables Continual Learning」論文では、中心的な式と理論セクション全体で逆 KL 発散を分析していますが、著者らが論文の結果を生み出すために使用したデフォルトの実装コードは、実際には順方向の KL を計算しています。
この不整合を発見した The logbook that caught it においても、著者自身のコードとデータを用いて検証を行った結果、論文で主張された +4pp という主要な成果を再現することはできませんでした。著者たちはすでに arXiv に明確化されたバージョンをアップロードしています。
評価がパディングによって希薄化されている
「トランスフォーマーには3つの射影が必要か?」という論文で、ある参加者が約 66% の評価対象ラベル位置が EOS パディングトークンであることを発見しました。これらのトークンは学習中にほぼゼロの損失に収束するため、パープレキシティ(モデルの不確実性)が約 3 分の 1 に低下してしまいます。この影響を補正すると、アブストラクトで主張されていた「キャッシュ削減率 50% に対する品質コストは 3.1%」という数値は、実際には約 9.4% に相当します。
🚨 誤った反証の発見
再現試行の中で欠陥が見つかるケースもありました。あるログブックでは「論文の手法はベースラインより 2 倍遅い」と断言していましたが、これは再現側の計算ミスでした。比較対象が「1 つのトラジェクトリあたりの時間」に対して「バッチサイズ 50 の場合の時間」を誤って用いていたのです。これを正しく正規化し直すと、参加者自身のデータは論文で主張されていた 8 倍の高速化を裏付ける結果となりました。
著者への連絡
確認された発見については、すべての著者に連絡を開始しました。アプローチはシンプルです。「我々が得た結論はこれであり、根拠となる証拠もすべて提示する。同意するか、それとも分析に誤りがあるか」。初期の反応は非常に好意的でした。
現時点では複数の論文で著者による発見の確認がなされ、2 つの arXiv 修正版が進行中です。またあるケースでは、著者がコンペティションの結果を知る約 1 ヶ月前に、新しい arXiv バージョンで誤りを自主的に修正していました。これは独立した収束としてカウントしました 🤗
人間の役割
このハッカソンが提起する最も興味深い問いは、「依然として人間による論文審査の役割はあるのか」という点です。我々は、いくつかの理由から「ある」と考えます。
エージェントの完全自律実行には現実的な限界がある。 エージェントは局所的なループに陥ったり、スケール依存の挙動を誤読したり(ページング論文に関するいくつかの「検証済み」判定は、ログ成長が可視化される前にチェックが停止した結果によるもの)、まれに単位不一致という前提の上に偽証を構築してしまうことさえあった。このチャレンジで最も信頼性の高い結果をもたらしたのは、人間が主導するワークフローだった。エージェントの方向転換や仮定の再考、あるいは計算リソースを1週間浪費する前に実験の前提自体が誤りだと判断することなどが含まれる。
現時点では、評価の一部は本質的に人間の介入を必要とする。 我々の 人間-in-the-loop 優勝者 が最も明確な例だ。論文は極端な量子化下での安定した画像生成を主張していた。数値指標では「崩壊なし」と示されたが、画像が実際に*使用可能か*という点は知覚的な問いだった。エージェントは専用のレビュー UI を構築し、人間が 128 ペアの画像すべてを手動で評価した。そのアノテーションはリポジトリにコミットされ、その後エージェントが一貫性を検証した。公開されたエージェントのトレースには、参加者がレビューツールの使い方を尋ね、「ペアを確認して CSV をリポジトリに置きましたのでご確認ください」と報告するまでのやり取りがすべて記録されている。
では、人間としてのレビュアーの役割とは何でしょうか。私たちはそれを「知能を効果的に管理すること」と捉えています。教授や研究責任者(PI)が、計算資源、ハルネス、データアクセス、そして適切なタイミングでのフィードバックを提供し、大学院生が良質な成果を出せる環境を整えるように、エージェントから最大の成果を引き出した参加者は、正しい環境を構築し、「適切な問い」を投げかけ、その後エージェントに実行を任せた人たちでした。
感謝の言葉
参加してくれた1,221名の方々に、受賞者、そして誠実な回答を寄せてくださった著者の皆様、さらにHugging FaceとalphaXivの運営チームに心から感謝いたします。このチャレンジで生成されたすべてのログブック、判定結果、トレース、およびアーティファクトは、チャレンジスペースから公開されています。これは現時点において、機械学習カンファレンスに対する最も大規模なオープンかつ主張ごとの監査であると考えています。ぜひこの記録が長く続くことは望んでいません。
今後の再現イベントにもご注目ください。🤗
原文を表示
- More papers than anyone can review
- The hackathon (July 15 - August 2nd)
- What we found Reproductions done well
- Falsifications, and what happened when we checked them
- Talking to authors
- The role of humans
- Thank you
Back in July, we ran a hackathon where more than 1,200 community members brought their own coding agents and tried to reproduce the papers published at ICML 2026, claim by claim. In 19 days, participants published 6,816 Trackio logbooks reproducing 2,226 papers, about a third of the conference 🤯
In this post, we're sharing what we learned from running this hackathon, and what it suggests about *the role humans will play* when agents are doing the research experiments.
More papers than anyone can review
Questions about how reproducible AI research really is are older than the current AI wave. But these questions are exacerbated by scale. ICML 2026 received 23,918 submissions and accepted 6,352 papers, roughly double the previous year, continuing an exponential trend that is at least partly driven by AI agents making it faster to run experiments and write them up.
Reviewing capacity has not doubled along with it. Reviewers at most conferences are volunteers who may not have the time or expertise to fully review a paper. Here is a review of one accepted ICML 2026 spotlight paper, in the reviewer's own words:
"My low confidence score is because I did not check all the proofs carefully."
Note that this paper got strong scores and a spotlight. Keep it in mind, because we will come back to this exact paper later in the post, and to what happened when we finally did check the proofs carefully.
What has changed, though, is that the same technology driving the flood of submissions can also help us keep up with it. Coding agents like Claude Code, Codex, Cursor, and Pi can now read a paper, write the code, launch the experiments, and report back on what they found. Checking a paper carefully used to cost a reviewer a weekend; an agent can attempt it in an afternoon, in parallel, thousands of times over.
So the question we wanted to ask was: if we actually re-examined a major conference at scale, and tried to reproduce every paper, what would we find?
The hackathon (July 15 - August 2nd)
Rather than audit papers ourselves, we opened it up to the whole community, with all the diversity of agent frameworks, compute budgets, and scientific taste that brings. From July 15 to August 2, 2026, the ICML 2026 Open Reproductions challenge worked like this:
- Pick a paper. We indexed all 6,341 accepted ICML 2026 papers with their abstracts and extracted the core scientific claims of each one, so an agent could start from a concrete, checkable target rather than a 40-page PDF. Multiple people reproducing the same paper was encouraged.
- Bring your own agent. Participants used Claude Code, Codex, Cursor, OpenResearch's orx, and everything in between. We provided a streamlined interface so an agent could pull the paper, its claims, and the challenge instructions with a single command.
- Reproduce, then publish everything. Every run produced a Trackio logbook: a static Hugging Face Space containing the write-up, the code that ran, the artifacts it produced, and (optionally) the full agent execution trace uploaded as a Hugging Face Dataset. The auditing process itself had to be auditable.
- Get judged. An automated Logbook Judge (running an open-weights model, GLM-5.2) re-read every logbook and issued a per-claim verdict: verified, falsified, toy (evidence at reduced scale), or inconclusive. The judge was explicitly instructed to treat each logbook's self-assessment as untrusted.
Participants received $20 in Hugging Face compute credits to run experiments on HF Jobs; across the challenge, participants launched 2,962 cloud jobs. Where a full reproduction was impossible, for example when a paper's dataset was proprietary or its checkpoints unreleased, participants ran toy reproductions on synthetic data mimicking the original's properties.
Here is what a finished reproduction looks like:
By the numbers, this hackathon was probably the largest attempted reproduction of a scientific conference:
- 1,221 community members joined the organization
- 6,816 reproduction logbooks published
- 2,226 papers attempted, 34% of the entire conference, many by several independent teams
- 35,908 claims judged, with all verdicts frozen in a public dataset at challenge close
- 2,962 HF Jobs launched; 274 full agent-trace datasets published on Hugging Face
What we found
Aggregating the claim-level verdicts per paper:
51% of examined papers (1,103) had at least one claim independently verified. Of those, 266 papers were fully reproduced, with every extracted claim verified, and 632 more were partially reproduced with nothing falsified. In total, 3,978 individual claims were confirmed with real experiments.
23% of examined papers (496) had at least one claim falsified or contested. That includes 49 papers where all claims were falsified and nothing could be verified, and, maybe most interestingly, 242 papers where independent reproduction teams reached *opposite* verdicts on the same claims. Reproducibility is not binary; it is adversarial.
The remainder sat in the middle: 502 papers with toy-scale evidence only, and 280 where nothing could be established either way (missing artifacts were the most common cause).
Reproductions done well
Some papers came through the gauntlet looking great, and the community's best logbooks are worth reading in their own right:
- "Flat Minima and Generalization: Insights from Stochastic Convex Optimization" was reproduced by 20 independent teams, 12 of which verified every claim. The one linked included and published the full agent trace.
- "A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness" had 14 of 17 logbooks verify every claim. A paper about unreliable LLM judges holding up under scrutiny by LLM agents :)
Falsifications, and what happened when we checked them
35 participants formally claimed they had falsified something. We adversarially re-verified every claimed falsification: re-reading the paper, re-reading the logbook, and re-deriving the math or re-implementing the experiment from the paper's own text.
A few of the confirmed falsifications, linking to the logbook that found it:
The paging paper from the introduction. The reviewer who did not check the proofs carefully? The paper, "Towards Optimal Robustness in Learning-Augmented Paging," claims its algorithm achieves robustness Hk+O(1)H_k + O(1). One participant's logbook measured the additive term growing like 0.38lnk0.38 \ln k and located the exact step of the proof that breaks. Our own re-implementation extended the sweep to k = 1,024 and confirmed the growth at roughly nine sigma. The true robustness is Hk+Θ(logk)H_k + \Theta(\log k).
A theorem that falls after step 224. "Attention's forward pass and Frank-Wolfe" proves that token particles collapse to the origin whenever the origin starts inside their convex hull. Three independent teams found counterexamples, with violations first appearing at t = 224, ~3,800, and 6,416 steps, which neatly explains why everyone else "verified" the claim: finite-horizon checks stop too early. The cleanest counterexample is stated in exact rational arithmetic, so there is no floating-point ambiguity to hide behind. The authors confirmed the same day and are working on a fix.
Theory written for one loss, results produced by another. In "Self-Distillation Enables Continual Learning," the paper's central equation and its entire theory section analyze reverse KL divergence, but the released code's default, which per the authors produced all the paper's results, computes forward KL. The logbook that caught it also failed to reproduce the paper's headline +4pp result under the authors' own code and data. The authors have already uploaded a clarified version to arXiv.
An evaluation diluted by padding. In "Do Transformers Need Three Projections?", a participant discovered that ~66% of evaluated label positions were EOS padding tokens that train to near-zero loss, deflating perplexity roughly threefold. The abstract's "3.1% quality cost for 50% cache reduction" becomes roughly 9.4% once corrected.
🚨 False falsifications. Sometimes we found flaws in an attempted reproduction. One logboook claimed dramatically, "the paper's method is 2x slower than the baseline"; this turned out to be an arithmetic bug in the *reproduction*: per-trajectory time compared against per-batch-of-50 time. Correctly normalized, the participant's own data confirms the paper's claimed 8x speedup.
Talking to authors
We have begun writing to the authors of every confirmed finding, with a simple framing: here is what we found, here is all the evidence, do you agree or is our analysis wrong? The early responses have been very positive:
So far authors have confirmed findings on multiple papers, two arXiv corrections are in flight, and in one case an author had quietly fixed the error in a new arXiv version a month before the challenge found it, which we count as independent convergence 🤗
The role of humans
The most interesting question that this hackathon raises is: do humans still have a role in reviewing papers? We think so, for several reasons:
Pure agent execution hits real limits. Agents got stuck in local loops, misread scale-dependent behavior (several "verified" verdicts on the paging paper came from checks that stopped before the log-k growth became visible), and occasionally built an entire falsification on top of a units mismatch. The challenge's most reliable results came from workflows where a human was steering: re-pointing the agent, questioning an assumption, or deciding that an experiment's premise was wrong before burning a week of compute on it.
Some evaluation is irreducibly human, for now. Our human-in-the-loop winner is the clearest example. The paper claimed stable image generation under extreme quantization. Numerical metrics said "no collapse"; whether the images were actually *usable* was a perceptual question. The agent built a purpose-built review UI, and the human personally judged all 128 image pairs, with the annotations committed to the repo and the agent validating their consistency afterward. The published agent trace captures the whole exchange, down to the participant asking how the review tool works and coming back with "I have gone over the pairs and put the csv in the repo, please check."
So what are our roles as human reviewers? We think it is to manage intelligence effectively. Much like a professor or principal investigator (PI) sets up an environment where grad students can do good work, with compute, harnesses, data access, and targeted feedback at the right moments, the participants who got the most out of their agents were the ones who built the right environment and *asked the right questions*, then let the agents do the running.
Thank you
To the 1,221 people who joined, the winners, the authors who responded with grace, and our organizers at Hugging Face and alphaXiv: thank you. Every logbook, verdict, trace, and artifact from the challenge is public, starting from the challenge Space. We think this is the largest open, claim-by-claim audit of a machine learning conference to date, and we would love for it not to hold that record for long.
Stay tuned for future reproduction events. 🤗
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み