AI スタートアップ Hark、自動操作用エージェント「Handoff」を発表
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
シリアルアントレプレナーのブレット・アッドコックが設立した AI スタートアップ Hark が、ウェブ操作における世界最高スコアを記録する自律型エージェント「Handoff」を発表し、一般登録を開始した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月7日 22:37
AI深層分析
キーポイント
Hark Handoff の発表と機能
シリアルアントレプレナーのブレット・アッドコックが設立した Hark が、ユーザーに代わってウェブを自律的に操作する「Handoff」を発表し、一般登録を開始した。
ベンチマークでの首位記録
同社によると、Handoff は人間評価型ベンチマーク「Online-Mind2Web (OM2W)」で 97.7 という史上最高スコアを記録し、OpenAI や Anthropic の主要モデルを上回った。
具体的な自律操作事例
Handoff は DoorDash での食事注文や United/Delta 航空の予約、LinkedIn での候補者へのメッセージ送信など、複雑なウェブタスクをエンドツーエンドで完遂できる。
リリーススケジュールとアクセス
一般登録は同日より開始され、ソフトウェアプラットフォームの初期リリースとして今月下旬から利用可能になる予定である。
競合モデルより大幅に低価格な料金設定
Hark Handoff は入力トークンあたり0.18ドル、出力トークンあたり2.37ドルで提供され、GPT-5.5の約10分の1のコストを実現している。
重要な引用
"among the top-performing in the world at navigating the open web on a user's behalf"
"Handoff recorded the top-ever score on Online-Mind2Web (OM2W)... posting a 97.7"
Hark also says it can serve the model at less than one-tenth the token price of competing frontier models — $0.18 per million input tokens and $2.37 per million output tokens, versus $5 and $30 for GPT 5.5 — with per-turn model latency of 0.8 seconds.
"is always working, it's looping"
編集コメントを表示
編集コメント
Hark が公開したベンチマークスコアは、大手 AI モデルとの明確な差を示しており、自律型エージェントの競争激化を象徴している。同社の技術が実際の複雑なタスクでどの程度安定して機能するか、今後の実運用での検証が注目される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
シークレットな AI スタートアップ「Hark」が、本日初の製品を発表しました。同社は今年初めに、連続起業家でありロボティクス研究者のブレット・アッドコック氏によって設立されました。
発表されたのは「Handoff」という名前の「コンピューター使用エージェント(CUA)」です。これは、ユーザーに代わってオープンウェブを自律的にナビゲートする能力において、世界最高峰の一つとされる製品です。具体的な機能としては、DoorDash での夕食注文や、United や Delta 航空での飛行機予約、LinkedIn での求人候補者へのメッセージ送信などが挙げられます。これらはすべて、エンドツーエンドで自律的に実行されます。
一般向けのサインアップは本日、hark.com で開始されました。同社のソフトウェアプラットフォームの初期リリースの一環として、今月下旬には利用可能になる予定です。
Handoff は、ウェブエージェントを対象とした第三者ベンチマーク「Online-Mind2Web (OM2W)」で、過去最高スコアを記録しました。このベンチマークは人間による評価に基づくリーダーボードを持つものです。Handoff のスコアは 97.7 で、OpenAI の GPT 5.4 が示した 92.8 を上回りました。また、Anthropic の Claude Opus 4.8(84.1)や Google の Gemini 2.5 Pro(69)も大きく引き離しています。

Hark はまた、競合する最先端モデルと比較してトークン価格の 10 分の 1 以下でモデルを提供できると述べています。具体的には、入力トークン 100 万あたり 0.18 ドル、出力トークン 100 万あたり 2.37 ドルです(GPT-5.5 はそれぞれ 5 ドルと 30 ドル)。また、1 トークンの処理遅延は 0.8 秒です。

各リクエストに対して、Handoff は専用のブラウザ、ファイルシステム、ターミナルを備えた仮想コンピューターを起動します。ユーザーは既存のアカウントを接続できるため、エージェントが保存された住所や支払い方法、履歴を使用してログインし、操作を行うことが可能になります。
Hark の調査によると、人々は毎日スクリーンタイムの 75% をブラウザで過ごしているにもかかわらず、公開 API を持つウェブサイトは 1000 に満たないのが現状です。このため、AI エージェントが業務を代行することは依然として課題となっています。
YouTube や SNS に投稿された約 4 分間の発表動画の中で、Adcock は倉庫の一角に座りながら(この空間は同社の成長過程を象徴するメタファーとしても機能しています)、Hark に対して口頭でリクエストを出します。「この場所をもっと賑やかにしよう……バラや桜を飾ってみようか」といった内容です。すると Handoff が花屋のウェブサイトへアクセスし、注文手続きを進める様子が映されます。Adcock はナレーションで、「一般的なチャットボットとは異なり、Handoff は常に稼働し、ループして動作する」と説明します。さらに彼は現在、採用活動の全工程を Handoff に任せていると語っています。
Hark の公式ブログ記事では、より多くのデモがリアルタイムおよび 5 倍速で公開されています。
しかし、Handoff については依然として大きな疑問が残っており、特に企業顧客や一般ユーザーにとって懸念材料となっています。
高得点のベンチマークだが、対象は前世代モデル
Hark が VentureBeat に提供した Handoff AI エージェントのベンチマーク比較データを見ると、対戦相手は GPT 5.5、GPT 5.4、Opus 4.8、Gemini 2.5 Pro です。これらはすべて前世代の最先端モデルです。
現在のリーダーである OpenAI の GPT-5.6 や Anthropic の Opus 5 は含まれていません。また、DeepSeek V4、Kimi K3、Qwen3.8-Max といった強力なオープンソースのコンピュータ操作対応モデルも除外されています。
これらの最新モデルは、Online-Mind2Web の結果をまだ公開しておらず、第三者もベンチマークの公式リーダーボードに投稿していません。そのため、Hark が主張する「史上最高」という記録は、現在利用可能な最強システムと比較することができません。
この省略は注目すべき点です。最新のフロンティアモデルが最も大きな進歩を示したのはまさにコンピュータ操作の分野であり、フルコンピュータ制御をカバーする関連ベンチマークである OSWorld 2.0 では、Anthropic の Opus 5 が約 70.6% を達成しています。一方、Hark が比較対象として選んだ Opus 4.8 モデルは 55.7% です。
レイテンシの比較にも同様の注意点があります。Hark が引用している GPT 5.5 と Opus 4.8 のターンあたり 6.8 秒、6 秒という数値は、競合モデルを最も高い(かつ最遅)推論レベルに設定した上で、Hark 独自の環境で測定されたものです。これを裏付ける独立したレイテンシ測定データは存在しません。
VentureBeat から Hark がこれらの最新モデルとの比較公開を検討しているか尋ねられた際、同社は具体的な回答をしていませんでした。
Hark が自ら選んだ比較対象内であっても、「最優位」という表現には注釈が必要です。WebTailBench v2 は Hark の結果表に含まれる 3 つのベンチマークの一つですが、GPT 5.5 は 72.3 を記録し、Handoff は 68.6 です。
3 つのベンチマークのうち 2 つ(WebTailBench と非公開の評価)は Hark 独自の環境で実行され、合格判定は同社内部の LLM ジャッジによって行われました。これは企業が完全にコントロールできる条件です。
Hark の価格競争力ははるかに明確です。Anthropic の最新モデル「Opus 5」は、前世代と同じく入力 100 万トークンあたり 5 ドル、出力 100 万トークンあたり 25 ドルのリスト価格を維持しています。したがって、Handoff が約 10 倍のコスト削減を実現している限り、そのベンチマーク性能が同等であれば、現在の最先端モデルに対してもこのコスト優位性は維持されると言えます。
トレーニングとファイルアクセス
Hark の研究プレビューでは、合理的なパイプラインが説明されています。発表前に VentureBeat に共有された資料によると、これは監督付き微調整(SFT)に続く非同期強化学習(GRPO アルゴリズム使用)という構成です。ただし同社は、現時点で実施したのはポストトレーニングのみであり、事前学習は「今年後半」に行う予定だと認めています。
つまり、Handoff は Hark が自ら訓練したベースモデルの上に構築されたものです。では、そのベースモデルは何なのか、また独自データとオープンデータのどの比率で訓練されたのかという問いに対し、Hark はまだ具体的な回答をしていません。
企業ユーザーにとって大きな疑問の一つは、専用仮想コンピュータおよびそこで作成されるファイルへのアクセス権限を誰が持つのかという点です。
Hark の広報担当者は、「セキュリティとプライバシーが最優先事項ですが、これは技術プレビュー段階である」と述べ、製品が夏末に市場投入された際にさらに詳細を共有すると付け加えました。
アドコック氏とハークの歩み
ハークは、アドコック氏が創業した4社目の企業です。彼はこの他にも、2018年に約1億ドルでアデッコグループに売却された人材マッチングプラットフォーム「Vettery」や、エアタクシーメーカーのアーチャー・アビエーション(Archer Aviation)、そしてヒューマノイドロボットユニコーンであるFigure AIを共同創業しました。
ハークは2026年5月、パークウェイ・ベンチャーキャピタルが主導し、Nvidia、AMD、Intel Capital、Qualcomm Ventures、Salesforce Ventures、ARK Investらが参加する形で、7億ドルのシリーズA資金調達を実施。企業価値は60億ドルに達しました。
同社の創設にはアドコック氏が自費で1億ドルを投入しており、Figureとハークの両社で同時に創業者兼CEOを務めています。これは広報担当者が確認した事実です。
両社の関係性について問われた際、広報担当者は「ハークのモデルはFigureのロボット上でトレーニングされているが、両社を統合する計画はない」と述べています。
アッドコック氏のプロモーションスタイルには懐疑的な声も上がっています。2025 年 4 月、フォーチュンのジェイソン・デル・レイ記者は、Figure が大々的に喧伝していた BMW との提携について報じました。その実態は、アッドコック氏が公言した「ロボット群がエンドツーエンドの運用を行う」という「フリート」の話とは遥かに控えめなものでした。BMW のスポークスマンであるスティーブ・ウィルソン氏は、単体の Figure ロボットが生産時間外に部品のピッキング練習を行っているだけだと明かしています。
しかし提携は進展し、2026 年 6 月現在、BMW は Figure 02 ロボットが 10 ヶ月の間に BMW X3 の生産に 3 万台以上貢献したと発表しました。また、次世代の Figure 03 ロボットも物流における部品シーケンシング(順序付け)の役割で同工場への導入が進められています。
ソーシャルネットワーク X では、アッドコック氏がこの報道を「誤った記述であり、明らかな嘘だ」と呼び、名誉毀損訴訟を起こすと脅しました。それから 2 ヶ月後、TechCrunch はアッドコック氏が技術カンファレンスで約束していたライブデモを欠席し、ステージ上で BMW 提携に関する質問に答えるのを避けたと報じました。
Handoff の数値が間違っているわけではない。このエージェントは非常に優秀であり、その価格設定(もし維持されれば)は主要な研究機関すべてを圧倒するものになるだろう。
原文を表示
Hark, the secretive AI startup founded earlier this yearby serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a "computer use agent" (CUA) that it says is among the top-performing in the world at navigating the open web on a user's behalf — ordering dinner on DoorDash, booking flights on United and Delta, or messaging job candidates on LinkedIn — all autonomously, end-to-end.
Sign-ups open to the public today at hark.com, with availability planned for later this month as part of the initial release of Hark's software platform.
The company says Handoff recorded the top-ever score on Online-Mind2Web (OM2W), a third-party benchmark with a human-evaluated leaderboard for web agents, posting a 97.7 against 92.8 for OpenAI's GPT 5.4, 84.1 for Anthropic's Claude Opus 4.8, and 69 for Google's Gemini 2.5 Pro.

Hark also says it can serve the model at less than one-tenth the token price of competing frontier models — $0.18 per million input tokens and $2.37 per million output tokens, versus $5 and $30 for GPT 5.5 — with per-turn model latency of 0.8 seconds.

For each request, Handoff spins up a dedicated virtual computer with its own browser, file system, and terminal, and users can connect existing accounts so the agent can log in and act with their saved addresses, payment methods, and history.
Hark's research uncovered that despite people spending 75% of their screentime every day in a browser, fewer than 1 in 1000 websites have publicly accessible APIs, making it challenging for AI agents to take over the workload.
In a roughly four-minute produced announcement video posted on YouTube and social media, Adcock — seated in a bare warehouse space that doubles as a metaphor for the company's build-out — speaks a request aloud to Hark ("let's liven this place up a bit… let's do some roses, maybe some cherry blossoms") and Handoff is shown navigating a florist's website to place the order, while Adcock narrates that unlike a typical chatbot, Handoff "is always working, it's looping," and says he now uses it for "all of my recruiting efforts end to end." In Hark's announcement blog post, more demos are shown in realtime and 5x speed.
But big some open questions about Handoff remain, especially for potential enterprise customers and users.
High-scoring benchmarks...but against last generation's models
Notably, the benchmark comparisons Hark provided to VentureBeat for its Handoff AI agent are against GPT 5.5, GPT 5.4, Opus 4.8, and Gemini 2.5 Pro — the prior generation of frontier models.
The current leaders, OpenAI's GPT-5.6 and Anthropic's Opus 5, are absent, as are strong open-source computer-use contenders like DeepSeek V4, Kimi K3, and Qwen3.8-Max.
These newer models haven't published Online-Mind2Web results, and no third party has posted them to the benchmark's public leaderboard— meaning Hark's "top-ever" claim cannot currently be checked against the strongest available systems.
The omission is notable because the newest frontier models have posted their largest gains precisely in computer use: onOSWorld 2.0, a related benchmark covering full computer control, Anthropic's Opus 5 scores roughly 70.6% versus 55.7% for the Opus 4.8 model Hark chose as its comparison point.
The latency comparison comes with similar caveats: the 6.8-second and 6-second per-turn figures Hark cites for GPT 5.5 and Opus 4.8 were measured by Hark, in Hark's own harness, with the competing models set to their highest — and slowest — reasoning level. No independent latency measurements exist for comparison.
Asked by VentureBeat whether Hark plans to publish comparisons against those newer models, the company did not specify.
Even within Hark's own chosen comparisons, the "best" framing has an asterisk: on WebTailBench v2, one of the three benchmarks in Hark's own results table, GPT 5.5 scores 72.3 to Handoff's 68.6.
Two of the three benchmarks (WebTailBench and an unnamed internal evaluation) were also run inside Hark's own harness, with pass rates computed by Hark's internal LLM judge — conditions the company controls.
Hark's pricing advantage is far clearer: Anthropic's newer Opus 5 carries the same $5-per-million-input and $25-per-million-output list price as its predecessor, so Handoff's roughly tenfold cost savings would hold up even against the current frontier — assuming its benchmark performance does too.
Training and file access
Hark's research preview describes a sensible-sounding pipeline — supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm, according to materials shared with VentureBeat prior to today's announcement — but the company acknowledges it has only done post-training so far, with pre-training "planned for later this year."
That means Handoff is built on top of a base model Hark did not train. Asked which base model it is, and what mix of proprietary and open data Handoff was trained on, Hark hasn't yet specified.
Another big question mark for enterprise users: who can access the dedicated virtual computers and the files created on them?
A Hark spokesperson said "security and privacy is a primary focus, but this is a technical preview," adding the company will share more when the product reaches market at the end of the summer.
Adcock's history leading up to Hark
Hark is Adcock's fourth company. He previously co-founded the talent marketplace Vettery (sold in 2018 for roughly $100 million), theair-taxi maker Archer Aviation, and the humanoid robotics unicorn Figure AI.
Hark raised a $700 million Series A round in May 2026 at a $6 billion valuation — led by Parkway Venture Capital, with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest.
Adcock seeded the company with $100 million of his own money and remains founder and CEO of both Figure and Hark simultaneously, a spokesperson confirmed.
Asked how the two companies interact, the spokesperson said Hark models "are being trained on the Figure robots," but that Adcock has no plans to combine them.
Adcock's promotional style has drawn skeptics. In April 2025, Fortune correspondentJason Del Rey reported that Figure's much-touted BMW partnership was far more modest than Adcock's public claims of a robot "fleet" performing "end-to-end operations": BMW spokesperson Steve Wilson said a single Figure robot was practicing picking up parts during non-production hours.
But the partnership has advanced, and as of June 2026, BMW said the Figure 02 robot supported production of more than 30,000 BMW X3 vehicles during a 10 month-period, and that the next-generation Figure 03 robot was being deployed at the plant for a parts-sequencing role in logistics.
On the social network X, Adcock called the story "mischaracterizations and downright lies"and threatened a defamation suit. Two months later,TechCrunch reported that Adcock skipped a promised live demo at a tech conference and sidestepped questions about the BMW deal onstage.
None of that means Handoff's numbers are wrong. The agent may well be excellent, and the pricing — if it holds — would undercut every major lab.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み