カウント・アンイシング(2 分読了):テキストガイド付き汎用オブジェクト計数モデルの提案
本文の状態
日本語全文を表示中
詳細モードで約3分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
TLDR AI
TLDR AI は、特定ドメインに依存せず多様な視覚領域や物体スケールに対応する、テキストガイド付きの一般化型オブジェクト計数モデル「Count Anything」を発表した。既存モデルの汎用性の低さを克服し、高精度な計数を可能にする。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
タイトル:Count Anything
**
要約:ドメイン固有のデータセットやタスク定義にまたがる物体数え上げは、一般化ビジョンモデルの急速な進展にもかかわらず依然として分断された状態にあります。既存の数え上げモデルは、群衆、車両、細胞、作物、またはリモートセンシングオブジェクトなどのシナリオ向けに特別に設計されていることが多く、カテゴリ間、視覚ドメイン間、物体スケール間、および密度分布間で一般化することに苦労します。本論文では、テキストガイド付きのドメイン横断的な物体数え上げを研究します。これは、モデルが画像と自然言語クエリを入力として受け取り、その基数(cardinality)が数え上げ結果となるインスタンス接地型のターゲット点セットを返すという設定です。この定式化は、カテゴリ条件付きの数え上げと解釈可能な空間位置特定を統合するものです。この設定をサポートするために、多様な公開データソースを統一されたベンチマークに再編成したクロスドメイン大規模物体数え上げデータセット「CLOC」を構築しました。CLOC は 6 つの視覚ドメイン(一般シーン、リモートセンシング、組織病理学、細胞顕微鏡、農業、微生物学)をカバーし、約 22 万枚の画像、619 カテゴリ、1500 万件の物体インスタンスを含みます。CLOC に基づき、テキストガイド付き物体数え上げ用の一般化モデル「Count Anything」を提案します。密度マップベースの方法(これが数え上げモデルを支配しています)とは異なり、「Count Anything」は離散的なインスタンス点を採用し、二重粒度のインスタンス列挙を実行します。「Region-level Sparse Counter(領域レベルのスプースカウンター)」は大規模でスパースなターゲットに対してオブジェクトレベルのアンカーを提供し、「Pixel-level Dense Counter(ピクセルレベルのデンスカウンター)」は密集した小さなターゲットや境界が不明瞭なターゲットを高密度点予測によって処理します。ポイント中心の教師あり学習戦略により、異種のアノテーションから学習が可能となり、相補的数え上げ融合(Complementary Count Fusion)により両方のカウンターをパラメータフリーな方法で組み合わせます。広範な実験により、「Count Anything」は高い精度と多ドメイン一般化能力を実現し、既存のオープンワールド数え上げ手法を上回ることを示しました。コードは以下の URL で利用可能です:this https URL。
**
主題:
コンピュータビジョンとパターン認識 (cs.CV)
引用形式:
arXiv:2605.30846 [cs.CV]
(またはこのバージョンについては
arXiv:2605.30846v1 [cs.CV])
https://doi.org/10.48550/arXiv.2605.30846
arXiv発行のDOI (DataCite経由)
## 提出履歴
From: Mengqi Lei [メールを表示]
[v1]**
2026年5月29日 (金) 05:08:31 UTC (41,518 KB)
原文を表示
Title:Count Anything
Abstract:Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehicles, cells, crops, or remote-sensing objects, and thus struggle to generalize across categories, visual domains, object scales, and density distributions. In this paper, we study text-guided object counting across domains, where a model takes an image and a natural-language query as input and returns an instance-grounded set of target points whose cardinality gives the count. This formulation unifies category-conditioned counting with interpretable spatial localization. To support this setting, we construct CLOC, a Cross-domain Large-scale Object Counting dataset that reorganizes diverse public data sources into a unified benchmark. CLOC covers six visual domains: General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, and Microbiology, with about 220K images, 619 categories, and 15M object instances. Based on CLOC, we propose Count Anything, a generalist model for text-guided object counting. Unlike density-map-based methods, which dominate counting models, Count Anything adopts discrete instance points and performs dual-granularity instance enumeration. A Region-level Sparse Counter provides object-level anchors for large and sparse targets, while a Pixel-level Dense Counter handles small, crowded, and weakly bounded targets via dense point prediction. A point-centric supervision strategy enables learning from heterogeneous annotations, and Complementary Count Fusion combines both counters in a parameter-free manner. Extensive experiments show that Count Anything achieves strong accuracy and multi-domain generalization, outperforming existing open-world counting methods. Code is available at: this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
|---|---|
| Cite as: | arXiv:2605.30846 [cs.CV] |
| (or arXiv:2605.30846v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.30846 arXiv-issued DOI via DataCite |
Submission history
From: Mengqi Lei [view email] [v1]
Fri, 29 May 2026 05:08:31 UTC (41,518 KB)
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み