AI スタートアップ Hark、自動操作用エージェント「Handoff」を発表
本文の状態
日本語全文を表示中
詳細モードで約9分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
VentureBeat AI
シークレット企業Harkが、ウェブ操作に特化したAIエージェント「Handoff」を発表し、既存の主要モデルを上回るベンチマークスコアと低コストを実現した。
AI深層分析を開く2026年8月6日 01:32
AI深層分析
キーポイント
高パフォーマンスなウェブ操作エージェントの登場
Harkは「Handoff」というAIエージェントを公開し、オンラインでの注文や予約などを自律的に実行する能力を持つことを発表した。
ベンチマークにおけるトップスコアと低コスト
同社によると、HandoffはWebエージェント向けベンチマークで97.7点を記録し、競合他社の最新モデルを上回るとしている。また、トークンあたりの価格も競合の10分の1以下であるという。
仮想環境による完全自律操作の実現
各リクエストごとに専用のブラウザやファイルシステムを持つ仮想コンピュータを起動し、ユーザーのアカウント情報を用いて実世界でのタスクを実行する仕組みを採用している。
ベンチマーク比較における懸念点
Harkが提示した比較対象は前世代モデルであり、現在の最上位モデルやオープンソースモデルとの直接比較データが存在しないため、実力評価には注意が必要である。
ベンチマークと評価手法の限界
Hark が提示した比較データは自社製のハルネスで測定されたものであり、競合モデルを最悪の設定でテストしているため独立検証が存在しない。
重要な引用
"is among the top-performing in the world at navigating the open web on a user's behalf"
"Handoff recorded the top-ever score on Online-Mind2Web (OM2W)... posting a 97.7 against 92.8 for OpenAI's GPT 5.4"
"serve the model at less than one-tenth the token price of competing frontier models"
"security and privacy is a primary focus, but this is a technical preview"
編集コメントを表示
編集コメント
Harkが提示したベンチマークスコアは前世代モデルとの比較であり、最新モデルやオープンソース系との対比がない点は留意が必要である。しかし、ウェブ操作特化型エージェントの低コスト化と自律性の向上という方向性は、実用化への大きな一歩となる可能性を秘めている。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
今年初にシリアルアントレプレナーかつロボティクスの専門家であるブレット・アコック氏によって設立された、謎めいた AI スタートアップ「Hark」が本日、「Handoff」という製品を発表しました。これは「コンピューター使用エージェント(CUA)」と呼ばれるもので、ユーザーの代わりにウェブをナビゲートする能力において世界最高レベルのパフォーマンスを発揮すると同社は主張しています。
具体的なタスクとしては、DoorDash での夕食注文や、United や Delta 航空での飛行機予約、LinkedIn での求職者へのメッセージ送信などが挙げられます。これらすべてが、エンドツーエンドで自律的に実行されます。
登録は本日 hark.com で一般公開され、今月下旬の Hark ソフトウェアプラットフォームの初期リリースに合わせて利用可能になる予定です。
同社によると、Handoff はウェブエージェント向けの第三者ベンチマークである「Online-Mind2Web(OM2W)」で過去最高スコアを記録しました。このベンチマークには人間による評価リーダーボードが用意されています。Handoff のスコアは 97.7 で、OpenAI の GPT-5.4 が 92.8、Anthropic の Claude Opus 4.8 が 84.1、Google の Gemini 2.5 Pro が 69 を記録した中で圧勝しました。
さらに Hark は、競合する最先端モデルと比較してトークン価格を 10 分の 1 以下に抑えることができるとしています。具体的には、入力トークン 100 万あたり 0.18 ドル、出力トークン 100 万あたり 2.37 ドルです(対照的に GPT-5.5 はそれぞれ 5 ドルと 30 ドル)。また、1 ターンあたりのモデルレイテンシは 0.8 秒です。
各リクエストに対して Handoff は、専用のブラウザ、ファイルシステム、ターミナルを備えた仮想コンピューターを起動します。ユーザーは既存のアカウントを接続できるため、エージェントは保存された住所や支払い方法、履歴情報を使用してログインし、行動することができます。
Hark の研究によると、人々は毎日スクリーンタイムの 75% をブラウザで過ごしているにもかかわらず、公開 API を持つウェブサイトは 1,000 サイトに満たないのが実情です。このため、AI エージェントが業務を代行するのは依然として困難な状況にあります。
YouTube や SNS に投稿された約 4 分間の発表動画では、Adcock が倉庫のような空間(同社の成長を象徴するメタファーとしても機能している)に座り、Hark に「この場所をもっと活気づけよう…バラや桜でも飾ろうか」と声をかけます。すると Handoff は花屋のウェブサイトを開き、注文手続きを完了します。Adcock はナレーションで、「一般的なチャットボットとは異なり、Handoff は常に稼働し、ループして動作する」と説明。さらに「採用活動の全工程に今では使っている」と語っています。Hark の公式ブログ記事では、リアルタイムおよび 5 倍速でのデモ映像も公開されています。
しかし、Handoff については依然として大きな疑問が残っており、特に企業顧客や一般ユーザーにとって懸念材料となっています。
高スコアなベンチマークだが、対象は前世代モデルのみ
特筆すべきは、Hark が VentureBeat に提供した Handoff AI エージェントのベンチマーク比較が、GPT 5.5、GPT 5.4、Opus 4.8、Gemini 2.5 Pro という、いずれも前世代の最先端モデルとの対比である点です。
現在のリーダー格である OpenAI の GPT-5.6 や Anthropic の Opus 5 はもちろん、DeepSeek V4、Kimi K3、Qwen3.8-Max といった有力なオープンソースのコンピュータ操作用エージェントも比較対象から除外されています。
これらの最新モデルは、Online-Mind2Web の結果をまだ公開しておらず、第三者がベンチマークの公開リーダーボードに投稿した記録もないため、Hark が主張する「史上最高」の性能は、現在利用可能な最強システムと比較して検証することができません。
この省略は注目すべき点です。最新のフロンティアモデルたちは、まさにコンピュータ操作において最大の進歩を遂げており、フルなコンピュータ制御をカバーする関連ベンチマークである OSWorld 2.0 では、Anthropic の Opus 5 が約 70.6% を達成しています。一方、Hark が比較対象として選んだ Opus 4.8 モデルは 55.7% です。
レイテンシの比較についても同様の注意点があります。Hark が引用する GPT-5.5 と Opus 4.8 の 1 ターンあたり 6.8 秒、6 秒という数値は、競合モデルを最も高い(そして最も遅い)推論レベルに設定した上で、Hark 独自のハッチで測定されたものです。これを裏付ける独立したレイテンシ測定データは存在しません。
VentureBeat から Hark がこれらの最新モデルとの比較公開を検討しているか問われた際、同社は具体的な回答を避けています。
Hark が自ら選んだ比較対象内であっても、「最優位」という表現には注釈が必要です。Hark の結果表に含まれる 3 つのベンチマークのうちの一つである WebTailBench v2 では、GPT-5.5 は 72.3 を記録し、Handoff は 68.6 です。
3 つのベンチマークのうち 2 つ(WebTailBench と名称不明の内部評価)は、Hark 独自のハッチ内で実行され、合格判定は Hark の内部 LLM ジャッジによって行われました。これは同社が完全に管理する条件です。
Hark の価格競争力は非常に明確です。Anthropic の最新モデル「Opus 5」は、入力 100 万トークンあたり 5 ドル、出力 100 万トークンあたり 25 ドルという従来モデルと同じリスト価格を維持しています。したがって、Handoff が約 10 倍のコスト削減を実現しているなら、そのベンチマーク性能も同等であれば、現在の最先端モデルに対してもこのコスト優位性は維持されることになります。
学習とファイルアクセスに関する仕組みについて、Hark の研究プレビューでは合理的なパイプラインが説明されています。発表前に VentureBeat に提供された資料によると、これは監督付き微調整(supervised fine-tuning)に続き、GRPO アルゴリズムを用いた非同期強化学習を行う構成です。しかし同社は現時点で学習の完了は後処理段階のみであり、事前学習(pre-training)については「今年後半に予定されている」と認めています。
つまり、Handoff は Hark が自ら訓練したベースモデルの上に構築されたものではなく、既存のモデルを基盤としています。では、そのベースモデルは何か、また独自データとオープンソースデータのどの比率で学習を行ったのかについて、Hark はいまだ具体的な回答をしていません。
企業ユーザーにとって大きな懸念材料となるのが、専用仮想コンピュータへのアクセス権限と、そこで作成されたファイルの管理責任です。誰がこれらのリソースを利用できるのでしょうか?
Hark の広報担当者は「セキュリティとプライバシーは最優先事項ですが、これは技術的なプレビュー段階です」と述べ、製品が夏末に市場投入される際に詳細を共有すると付け加えました。
Adcock 氏による Hark 設立までの経緯について
Hark は Adcock 氏が率いる 4 つ目の企業です。彼は以前、2018 年に約 1 億ドルで売却された人材マッチングプラットフォーム「Vettery」の共同創設者であり、エアタクシーメーカーの「Archer Aviation」、そしてヒューマノイドロボットユニコーンの「Figure AI」でもリーダーシップを発揮しました。
Hark は 2026 年 5 月、シリーズ A ラウンドで 7 億ドルを調達し、企業価値は 60 億ドルに達しました。このラウンドは Parkway Venture Capital が主導し、Nvidia、AMD、Intel Capital、Qualcomm Ventures、Salesforce Ventures、ARK Invest が参加しています。
Adcock は創業時に自らの資金 1 億ドルを投入しており、Figure と Hark の両社の創業者兼 CEO を兼任し続けています。これは広報担当者が確認した事実です。
両社の関係性について問われた際、広報担当者は「Hark のモデルは Figure ロボット上でトレーニングされているが、Adcock 氏には両社を統合する計画はない」と述べています。
Adcock 氏のプロモーション手法には懐疑的な声も上がっています。2025 年 4 月、Fortune のジェイソン・デル・レイ記者は、Figure が大々的に宣伝していた BMW との提携が、ロボットによる「エンドツーエンドの運用」を行う「フリート(隊列)」を構築したという Adcock 氏の公言ほど大規模なものではないと報じました。BMW の広報担当者スティーブ・ウィルソン氏は、単一の Figure ロボットが生産時間外に部品のピッキング練習を行っているだけだと明かしています。
しかし提携は進展しており、2026 年 6 月現在、BMW は Figure 02 ロボットが 10 ヶ月の期間中に 3 万 3000 台以上の BMW X3 の生産をサポートしたと発表しました。また、次世代の Figure 03 ロボットは物流における部品シーケンシング業務のために同工場へ導入されつつあることも明らかにしています。
SNS「X」上で Adcock はこの報道を「誤った記述であり、まさに嘘だ」と断じ、名誉毀損訴訟を起こすと脅しました。それから 2 ヶ月後、TechCrunch は Adcock 氏が技術カンファレンスで約束していたライブデモを欠席し、ステージ上で BMW 提携に関する質問に答えるのを避けたと報じています。
Handoff の数値が間違っているわけではない。このエージェントは非常に優秀であり、もし価格設定が維持されれば、主要な研究機関のすべてを圧倒する安さになるだろう。
原文を表示
Hark, the secretive AI startup founded earlier this year by serial entrepreneur and roboticist Brett Adcock, today announced Handoff, a "computer use agent" (CUA) that it says is among the top-performing in the world at navigating the open web on a user's behalf — ordering dinner on DoorDash, booking flights on United and Delta, or messaging job candidates on LinkedIn — all autonomously, end-to-end.
Sign-ups open to the public today at hark.com, with availability planned for later this month as part of the initial release of Hark's software platform.
The company says Handoff recorded the top-ever score on Online-Mind2Web (OM2W), a third-party benchmark with a human-evaluated leaderboard for web agents, posting a 97.7 against 92.8 for OpenAI's GPT 5.4, 84.1 for Anthropic's Claude Opus 4.8, and 69 for Google's Gemini 2.5 Pro.
Hark also says it can serve the model at less than one-tenth the token price of competing frontier models — $0.18 per million input tokens and $2.37 per million output tokens, versus $5 and $30 for GPT 5.5 — with per-turn model latency of 0.8 seconds.
For each request, Handoff spins up a dedicated virtual computer with its own browser, file system, and terminal, and users can connect existing accounts so the agent can log in and act with their saved addresses, payment methods, and history.
Hark's research uncovered that despite people spending 75% of their screentime every day in a browser, fewer than 1 in 1000 websites have publicly accessible APIs, making it challenging for AI agents to take over the workload.
In a roughly four-minute produced announcement video posted on YouTube and social media, Adcock — seated in a bare warehouse space that doubles as a metaphor for the company's build-out — speaks a request aloud to Hark ("let's liven this place up a bit… let's do some roses, maybe some cherry blossoms") and Handoff is shown navigating a florist's website to place the order, while Adcock narrates that unlike a typical chatbot, Handoff "is always working, it's looping," and says he now uses it for "all of my recruiting efforts end to end." In Hark's announcement blog post, more demos are shown in realtime and 5x speed.
But big some open questions about Handoff remain, especially for potential enterprise customers and users.
High-scoring benchmarks...but against last generation's models
Notably, the benchmark comparisons Hark provided to VentureBeat for its Handoff AI agent are against GPT 5.5, GPT 5.4, Opus 4.8, and Gemini 2.5 Pro — the prior generation of frontier models.
The current leaders, OpenAI's GPT-5.6 and Anthropic's Opus 5, are absent, as are strong open-source computer-use contenders like DeepSeek V4, Kimi K3, and Qwen3.8-Max.
These newer models haven't published Online-Mind2Web results, and no third party has posted them to the benchmark's public leaderboard — meaning Hark's "top-ever" claim cannot currently be checked against the strongest available systems.
The omission is notable because the newest frontier models have posted their largest gains precisely in computer use: on OSWorld 2.0, a related benchmark covering full computer control, Anthropic's Opus 5 scores roughly 70.6% versus 55.7% for the Opus 4.8 model Hark chose as its comparison point.
The latency comparison comes with similar caveats: the 6.8-second and 6-second per-turn figures Hark cites for GPT 5.5 and Opus 4.8 were measured by Hark, in Hark's own harness, with the competing models set to their highest — and slowest — reasoning level. No independent latency measurements exist for comparison.
Asked by VentureBeat whether Hark plans to publish comparisons against those newer models, the company did not specify.
Even within Hark's own chosen comparisons, the "best" framing has an asterisk: on WebTailBench v2, one of the three benchmarks in Hark's own results table, GPT 5.5 scores 72.3 to Handoff's 68.6.
Two of the three benchmarks (WebTailBench and an unnamed internal evaluation) were also run inside Hark's own harness, with pass rates computed by Hark's internal LLM judge — conditions the company controls.
Hark's pricing advantage is far clearer: Anthropic's newer Opus 5 carries the same $5-per-million-input and $25-per-million-output list price as its predecessor, so Handoff's roughly tenfold cost savings would hold up even against the current frontier — assuming its benchmark performance does too.
Training and file access
Hark's research preview describes a sensible-sounding pipeline — supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm, according to materials shared with VentureBeat prior to today's announcement — but the company acknowledges it has only done post-training so far, with pre-training "planned for later this year."
That means Handoff is built on top of a base model Hark did not train. Asked which base model it is, and what mix of proprietary and open data Handoff was trained on, Hark hasn't yet specified.
Another big question mark for enterprise users: who can access the dedicated virtual computers and the files created on them?
A Hark spokesperson said "security and privacy is a primary focus, but this is a technical preview," adding the company will share more when the product reaches market at the end of the summer.
Adcock's history leading up to Hark
Hark is Adcock's fourth company. He previously co-founded the talent marketplace Vettery (sold in 2018 for roughly $100 million), the air-taxi maker Archer Aviation, and the humanoid robotics unicorn Figure AI.
Hark raised a $700 million Series A round in May 2026 at a $6 billion valuation — led by Parkway Venture Capital, with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest.
Adcock seeded the company with $100 million of his own money and remains founder and CEO of both Figure and Hark simultaneously, a spokesperson confirmed.
Asked how the two companies interact, the spokesperson said Hark models "are being trained on the Figure robots," but that Adcock has no plans to combine them.
Adcock's promotional style has drawn skeptics. In April 2025, Fortune correspondent Jason Del Rey reported that Figure's much-touted BMW partnership was far more modest than Adcock's public claims of a robot "fleet" performing "end-to-end operations": BMW spokesperson Steve Wilson said a single Figure robot was practicing picking up parts during non-production hours.
But the partnership has advanced, and as of June 2026, BMW said the Figure 02 robot supported production of more than 30,000 BMW X3 vehicles during a 10 month-period, and that the next-generation Figure 03 robot was being deployed at the plant for a parts-sequencing role in logistics.
On the social network X, Adcock called the story "mischaracterizations and downright lies" and threatened a defamation suit. Two months later, TechCrunch reported that Adcock skipped a promised live demo at a tech conference and sidestepped questions about the BMW deal onstage.
None of that means Handoff's numbers are wrong. The agent may well be excellent, and the pricing — if it holds — would undercut every major lab.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み