Hugging Face、エッジ向けに高速・高品質な視覚機能を持つ LFM2.5-VL-3B を公開
本文の状態
日本語全文を表示中
詳細モードで約11分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
Hugging Face Blog
Liquid AI は Hugging Face Blog で、エッジデバイス向けに設計された軽量ビジョン言語モデル「LFM2.5-VL-3B」の公開を発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月12日 23:56
AI深層分析
キーポイント
エッジ特化型モデルの発表
Liquid AI が、エッジデバイスでの推論速度と能力を向上させることを目的とした「LFM2.5-VL-3B」モデルを発表した。
軽量かつ高性能な設計
30 億パラメータという軽量サイズでありながら、視覚的能力の向上と高速処理を両立させる技術的革新が示された。
Hugging Face での公開
同モデルは Hugging Face Blog を通じて公開され、開発者がすぐにアクセスして利用可能な状態にある。
エッジデバイス向けの高機能ビジョン言語モデル
LFM2.5-VL-3Bは独自ハードウェアで実行可能な最も高性能なビジョン言語モデルであり、ドキュメントや画面の理解、物体の特定、ツール呼び出しに対応する。即時応答を重視し推論プロセスを省略することで、リアルタイムおよびオンデバイスアプリでの高速動作を実現している。
4つの主要機能強化
本モデルは画面/UI理解の向上、自然言語による物体検出とグラウンディングの改善、複数画像間の推論能力強化、およびテキスト・ビジョン両方の状況下での関数呼び出し性能の大幅な強化を特徴とする。
重要な引用
LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Back to Articles
It understands documents and screens alike, grounds objects, and can call tools.
It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps.
編集コメントを表示
編集コメント
エッジデバイスにおける推論速度と精度の両立は長年の課題であり、3B という軽量サイズでの実現は実用化への大きな一歩となる。開発者はこのモデルをベースに、特定のユースケース向けのカスタマイズや最適化を進めることができるだろう。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
LFM2.5-VL-3B は、独自のハードウェアで実行可能な最も高性能なビジョン・ランゲージモデルです。文書や画面の理解、物体の特定(グラウンディング)、ツール呼び出しが可能で、推論プロセスを経ずに直接回答するため、リアルタイム処理やオンデバイスアプリでも高速な応答を実現します。
LFM2.5-VL-3B は、以前のリリースに比べて以下の 4 つの主要な改善点により、ビジョン・ランゲージ能力を拡張しています:
- 画面/UI の理解: 異なるデバイスのデジタル画面に対する強力な理解力。
- グラウンディング: 自然言語によるクエリに対応した、精度の高い物体検出と位置特定機能の向上。
- 複数画像入力: 複数の画像にわたる推論能力の強化。
- 関数呼び出し(Function calling): テキストのみ、およびビジョン・テキストの両方の状況において、関数呼び出し能力が大幅に強化されました。
最も高性能なビジョン・ランゲージモデルの開発プロセス
LFM2.5-VL-3B は、SigLIP2 400M NaFlex ビジョンエンコーダー と、テキストモデル「LFM2.5-2.6B」と同じ事前学習済みバックボーンを組み合わせて構築されています。約 34T トークンで事前学習を行い、そのうちビジョンデータは従来比 4 倍に増やしました。学習データには、厳選された合成画像キャプション、OCR(光学文字認識)、グラウンディング、指示追従のセットが含まれています。
ラテン文字以外のスクリプトにも対応するため、ゼロから再トレーニングするのではなく、トークナイザーをその場で拡張することで、語彙サイズを 128K に倍増させました。詳細はこちら。
ポストトレーニングは 2 つの段階で実行されます。第一段階は、大規模な教師モデルからの知識蒸留と Antidoom training を用いた教師あり微調整(SFT)です。第二段階は、多様な報酬を用いた強化学習(RL)となります。
ベンチマーク結果
LFM2.5-VL-3B の性能を、画像認識とテキスト処理の両方のベンチマークで評価しました。
画像認識ベンチマークでは、多言語での視覚的理解、指示への従順さ、数学・科学推論、ドキュメント理解、物体検出、複数画像の統合的理解、そして画面操作の理解を網羅しています。LFM2.5-VL-3B は実世界の写真タスクにおいて同サイズクラスでトップの性能を発揮する一方、文書やチャート、画面上の UI 要素といったデジタルコンテンツの読み取りにおいても優れた能力を示します。
| タスク | ベンチマーク | LFM2.5-VL-3B (3.1B) | LFM-VL-3B (3.1B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | InternVL 3.5 2B (2.4B) | InternVL 3.5 4B (4.7B) | Qwen3.5-2B (2.3B) | Qwen3.5-4B (4.7B) |
|---|---|---|---|---|---|---|---|---|---|
| 一般能力 | MMStar | 63.3 | 57.7 | 45.3 | 52.9 | 57.7 | 65.5 | 55.1 | 59.3 |
| 一般能力 | MME | 73.1 | 73.0 | 54.9 | 67.6 | 73.6 | 81.0 | 76.2 | 79.5 |
| 一般能力 | RealWorldQA | 73.1 | 71.1 | 60.0 | 64.3 | 61.6 | 67.7 | 65.1 | 67.1 |
| 一般能力 | SimpleVQA | 35.4 | 33.0 | 27.3 | 30.4 | 30.5 | 33.7 | 35.2 | 40.7 |
| 一般能力 | SEED-Bench (image) | 77.7 | 76.6 | 71.4 | 75.3 | 75.4 | 76.4 | 75.8 | 76.1 |
| 一般能力 | MMBench (dev EN v1.1) | 81.0 | 80.0 | 64.2 | 71.6 | 76.2 | 81.1 | 73.1 | 78.4 |
| 一般能力 | CountBenchQA | 87.3 | 92.2 | 70.4 | 80.5 | 70.4 | 82.5 | 83.8 | 86.7 |
| 多言語 | MMMB | 83.0 | 81.9 | 73.3 | 80.4 | 76.3 | 81.5 | 75.9 | 82.0 |
| 多言語 | Multilingual MMBench | 79.5 | 76.3 | 62.8 | 71.2 | 70.9 | 76.6 | 69.9 | 77.0 |
| マルチモーダル指示従順性 | MM-IFEval | 60.6 | 51.4 | 65.6 | 68.2 | 47.1 | 54.5 | 55.4 | 63.1 |
| STEM | LogicVista | 37.4 | 32.2 | 29.5 | 34.5 | 30.9 | 36.2 | 34.0 | 37.6 |
| STEM | MathVista (mini) | 68.5 | 62.1 | 37.8 | 45.2 | 56.8 | 67.1 | 48.7 | 63.6 |
| STEM | MMMU-Pro | 30.5 | 28.7 | 26.9 | 32.6 | 21.3 | 22.7 | 24.9 | 36.0 |
| STEM | MMMU (val) | 48.4 | 45.6 | 41.1 | 49.3 | 52.0 | 60.7 | 44.1 | 50.3 |
| ドキュメント、OCR & チャート | ChartQA (test) | 81.3 | 80.4 | 43.2 | 42.1 | 81.7 | 86.2 | 78.4 | 84.2 |
| ドキュメント、OCR & チャート | DocVQA (val) | 91.1 | 89.8 | 85.7 | 87.4 | 88.4 | 91.8 | 92.6 | 94.8 |
| ドキュメント、OCR & チャート | InfographicVQA (val) | 70.2 | 67.8 | 54.4 | 60.9 | 69.3 | 76.9 | 73.5 | 80.3 |
| ドキュメント、OCR & チャート | OCRBench v1 | 84.2 | 81.7 | 70.2 | 73.5 | 83.9 | 82.0 | 84.4 | 85.6 |
| ドキュメント、OCR & チャート | OCRBench v2 (En) | 47.5 | 43.9 | 44.4 | 48.8 | 45.5 | 49.1 | 47.7 | 58.7 |
| ドキュメント、OCR & チャート | TextVQA (val) | 84.3 | 83.0 | 62.5 | 69.0 | 76.6 | 77.5 | 77.3 | 81.2 |
| グラウンディング | RefCOCO-avg | 87.9 | 57.1 | 67.3 | 72.1 | 82.9 | 88.8 | 78.5 | 86.6 |
| マルチ画像 | BLINK | 61.5 | 50.2 | 45.2 | 52.2 | 52.0 | 57.2 | 48.6 | 58.7 |
| マルチ画像 | MuirBench | 58.3 | 34.9 | 32.9 | 51.8 | 45.0 | 53.5 | 48.2 | 62.0 |
| ハルシネーション(幻覚) | HallusionBench | 47.2 | 46.4 | 41.8 | 49.8 | 47.6 | 52.1 | 49.3 | 51.7 |
| ハルシネーション(幻覚) | POPE | 88.7 | 89.2 | 84.0 | 86.9 | 88.0 | 88.9 | 88.6 | 86.0 |
| GUI | ScreenSpot-v2 Desktop | 78.7 | 6.0 | 28.1 | 45.8 | 79.9 | 82.0 | 63.8 | 76.3 |
| GUI | ScreenSpot-v2 Mobile | 81.2 | 7.6 | 42.9 | 60.3 | 86.2 | 87.8 | 69.7 | 81.4 |
| GUI | ScreenSpot-v2 Web | 82.2 | 2.5 | 22.4 | 47.6 | 79.9 | 82.6 | 65.9 | 77.8 |
| 平均 | - | 69.4 | 57.2 | 52.0 | 59.7 | 64.6 | 69.4 | 63.7 | 70.1 |
表内のすべての数値は 0〜100 に正規化されています。評価には vLLM 0.26.0 を使用し、各モデルの推奨生成パラメータ(利用可能な場合)を適用しました。すべてのテストで推論モードはオフとし、モデルに推論過程を経ずに直接回答するようプロンプトしています。
また、LFM2.5-VL-3B の指示従順性とツール使用能力についても、テキストのみを用いたベンチマークで評価を行いました。その結果、指示従順性は全体的に向上し、特にツール使用能力では劇的な改善が見られました。ツール使用においては、Gemma-4-E2B や Qwen3.5-2B と同等の性能を発揮しています。
| タスク | ベンチマーク | LFM2.5-VL-3B (3.1B) | LFM-VL-3B (3.1B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | InternVL 3.5 2B (2.4B) | InternVL 3.5 4B (4.7B) | Qwen3.5-2B (2.3B) | Qwen3.5-4B (4.7B) |
|---|---|---|---|---|---|---|---|---|---|
| 指示従順性 | IFEval | 82.3 | 72.9 | 83.0 | 87.9 | 32.4 | 35.4 | 73.6 | 86.2 |
| 指示従順性 | IFBench | 25.8 | 20.8 | 34.1 | 39.2 | 24.4 | 24.5 | 28.9 | 33.5 |
| 指示従順性 | Multi-IF | 59.4 | 46.5 | 69.4 | 77.4 | 16.3 | 16.9 | 53.5 | 66.7 |
| ツール使用・関数呼び出し | ToolSandbox | 59.5 | 26.4 | 56.5 | 61.6 | N/A | N/A | 47.7 | 65.0 |
| ツール使用・関数呼び出し | BFCL V4 | 32.5 | 20.5 | 33.2 | 40.0 | N/A | N/A | 33.9 | 53.6 |
InternVL 3.5 モデルは関数呼び出しには対応していません。
この結果から、LFM2.5-VL-3B は汎用的なビジョン・ランゲージモデルとして非常に強力であることがわかります。キャプション生成や視覚的質問応答(VQA)、ドキュメント理解といった日常的なタスクをカバーしており、特に物体の特定(グラウンディング)、画面や文書の読み取り、ツール呼び出しにおいて優れた性能を発揮します。
CPU および GPU での推論速度
LFM2.5-VL-3B は、llama.cpp、MLX、vLLM、SGLang、ONNX など、推論エコシステム全体でリリース初日からサポートされています。
オンデバイス推論。 LFM2.5-VL-3B は M5 Max で 1 秒あたり 228 トークン、Ryzen AI Max+ 395 で 116 トークンをデコードします。メモリ使用量は約 3GB に収まります。さらに Galaxy S26 Ultra では 1 秒あたり 20 トークンの速度を達成するため、端末内で完結して実行することも可能です。
GPU 推論。 LFM2.5-VL-3B は遅延を常に低く抑え、マルチフレーム入力においても最速の性能を示します。LFM2.5-VL-3B は、テストしたすべてのモデルの中で出力スループットも最も高く、高同時接続時で 1 秒あたり約 11K トークンを達成しました。これはより大きな 4B クラスのモデルの約 2 倍であり、小さな 2B クラスのモデルをも上回ります。単一の H100 GPU であれば、1 日あたりの出力トークン数は約 10 億に達します。
LFM2.5-VL-3B の使い方
大量のワークロードでオンデバイスでの知能処理が必要な場合は、LFM2.5-VL-3B をご活用ください。
transformers の最新バージョン(transformers>=5.0.0 に互換性あり)をインストールしてください:
%pip install -q torch torchvision accelerate "transformers>=5.10.1"
その後、モデルを読み込んで実行します:
import torch
from transformers.image_utils import load_image
from transformers import AutoModelForImageTextToText, AutoProcessor
from IPython.display import display
MODEL_ID = "LiquidAI/LFM2.5-VL-3B"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForImageTextToText.from_pretrained(
MODEL_ID,
device_map="auto",
dtype="bfloat16",
)
img_url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/coco_sample.png"
input_image = load_image(img_url)
display(input_image)
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": input_image},
{"type": "text", "text": "Describe this image in two concise sentences."},
],
}
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
do_sample=True,
temperature=0.2,
top_k=50,
repetition_penalty=1.0,
max_new_tokens=256,
)
output = processor.batch_decode(outputs[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0]
print(output)
Two cats are sleeping on a pink couch with two remote controls.
LFM2.5-VL-3B の多画像入力、グラウンディング、OCR、ツール呼び出しなどへの活用方法は、公式ドキュメント で詳しく紹介しています。動画のサンプルについては、リリースブログ をご覧ください。
LFM2.5-VL-3B デモ
この ブラウザデモ では、LFM2.5-VL-3B が搭載した視覚対応チャットインターフェースを実際に体験できます。複数の画像を撮影またはアップロードし、グラウンディングや OCR、ツール利用を通じてモデルと対話することが可能です。
始め方
LFM2.5-VL-3B は本日より Hugging Face で公開されています。
「AI をどこでも動かす」という私たちのビジョンを実現するため、以下のモデルを提供しています。
- ダウンロード: Hugging Face の LFM2.5-VL-3B
- 試す: 設定不要で ブラウザの WebGPU デモ を実行可能
- ファインチューニング: チュートリアル に従って、LFM2.5-VL-3B をあなたのタスクに合わせて調整できます。
みなさんの作品が楽しみです。
引用について
本記事は以下のように引用してください:
Liquid AI, "LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge", Liquid AI Blog, Aug 2026.
または BibTeX 形式で引用することも可能です:
@article{liquidAI2026VL3B,
author = {Liquid AI},
title = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-vl-3b},
}
原文を表示
LFM2.5-VL-3B is our most capable vision-language model you can run on your own hardware. It understands documents and screens alike, grounds objects, and can call tools. It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps.
LFM2.5-VL-3B extends the vision-language capabilities of our previous releases with four major improvements:
- Screen/UI understanding: Strong understanding of digital screens across different devices.
- Grounding: Improved grounding and object detection with natural language queries.
- Multi-image input: Improved reasoning across multiple images.
- Function calling: Significantly stronger at function calling, in text-only and vision-text situations.
How we trained our most capable vision-language model
LFM2.5-VL-3B pairs a SigLIP2 400M NaFlex vision encoder with the same pre-trained backbone as our LFM2.5-2.6B text model. It is pre-trained on about 34T tokens, with 4x more vision data than before, drawn from curated and synthetic image-caption, OCR, grounding, and instruction-following sets. To support non-Latin scripts, we doubled the vocabulary to 128K by extending the tokenizer in place rather than retraining from scratch.
Post-training runs in two stages: First is supervised fine-tuning (SFT), with knowledge distillation from a larger teacher and Antidoom training. Second is multi-reward reinforcement learning (RL).
Benchmark results
We evaluated LFM2.5-VL-3B across both vision and text benchmarks.
The vision benchmarks cover multilingual visual comprehension, instruction following, visual math and scientific reasoning, document understanding, object detection, multi-image understanding, and screen understanding. LFM2.5-VL-3B leads its size class on real-world image tasks, while also reading digital content well, from documents and charts to on-screen UI elements.
| Task | Benchmark | LFM2.5-VL-3B (3.1B) | LFM2-VL-3B (3.1B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | InternVL 3.5 2B (2.4B) | InternVL 3.5 4B (4.7B) | Qwen3.5-2B (2.3B) | Qwen3.5-4B (4.7B) |
|---|---|---|---|---|---|---|---|---|---|
| General | MMStar | 63.3 | 57.7 | 45.3 | 52.9 | 57.7 | 65.5 | 55.1 | 59.3 |
| General | MME | 73.1 | 73.0 | 54.9 | 67.6 | 73.6 | 81.0 | 76.2 | 79.5 |
| General | RealWorldQA | 73.1 | 71.1 | 60.0 | 64.3 | 61.6 | 67.7 | 65.1 | 67.1 |
| General | SimpleVQA | 35.4 | 33.0 | 27.3 | 30.4 | 30.5 | 33.7 | 35.2 | 40.7 |
| General | SEED-Bench (image) | 77.7 | 76.6 | 71.4 | 75.3 | 75.4 | 76.4 | 75.8 | 76.1 |
| General | MMBench (dev EN v1.1) | 81.0 | 80.0 | 64.2 | 71.6 | 76.2 | 81.1 | 73.1 | 78.4 |
| General | CountBenchQA | 87.3 | 92.2 | 70.4 | 80.5 | 70.4 | 82.5 | 83.8 | 86.7 |
| Multilingual | MMMB | 83.0 | 81.9 | 73.3 | 80.4 | 76.3 | 81.5 | 75.9 | 82.0 |
| Multilingual | Multilingual MMBench | 79.5 | 76.3 | 62.8 | 71.2 | 70.9 | 76.6 | 69.9 | 77.0 |
| Multimodal IF | MM-IFEval | 60.6 | 51.4 | 65.6 | 68.2 | 47.1 | 54.5 | 55.4 | 63.1 |
| STEM | LogicVista | 37.4 | 32.2 | 29.5 | 34.5 | 30.9 | 36.2 | 34.0 | 37.6 |
| STEM | MathVista (mini) | 68.5 | 62.1 | 37.8 | 45.2 | 56.8 | 67.1 | 48.7 | 63.6 |
| STEM | MMMU-Pro | 30.5 | 28.7 | 26.9 | 32.6 | 21.3 | 22.7 | 24.9 | 36.0 |
| STEM | MMMU (val) | 48.4 | 45.6 | 41.1 | 49.3 | 52.0 | 60.7 | 44.1 | 50.3 |
| Document, OCR & Chart | ChartQA (test) | 81.3 | 80.4 | 43.2 | 42.1 | 81.7 | 86.2 | 78.4 | 84.2 |
| Document, OCR & Chart | DocVQA (val) | 91.1 | 89.8 | 85.7 | 87.4 | 88.4 | 91.8 | 92.6 | 94.8 |
| Document, OCR & Chart | InfographicVQA (val) | 70.2 | 67.8 | 54.4 | 60.9 | 69.3 | 76.9 | 73.5 | 80.3 |
| Document, OCR & Chart | OCRBench v1 | 84.2 | 81.7 | 70.2 | 73.5 | 83.9 | 82.0 | 84.4 | 85.6 |
| Document, OCR & Chart | OCRBench v2 (En) | 47.5 | 43.9 | 44.4 | 48.8 | 45.5 | 49.1 | 47.7 | 58.7 |
| Document, OCR & Chart | TextVQA (val) | 84.3 | 83.0 | 62.5 | 69.0 | 76.6 | 77.5 | 77.3 | 81.2 |
| Grounding | RefCOCO-avg | 87.9 | 57.1 | 67.3 | 72.1 | 82.9 | 88.8 | 78.5 | 86.6 |
| Multi-Image | BLINK | 61.5 | 50.2 | 45.2 | 52.2 | 52.0 | 57.2 | 48.6 | 58.7 |
| Multi-Image | MuirBench | 58.3 | 34.9 | 32.9 | 51.8 | 45.0 | 53.5 | 48.2 | 62.0 |
| Hallucination | HallusionBench | 47.2 | 46.4 | 41.8 | 49.8 | 47.6 | 52.1 | 49.3 | 51.7 |
| Hallucination | POPE | 88.7 | 89.2 | 84.0 | 86.9 | 88.0 | 88.9 | 88.6 | 86.0 |
| GUI | ScreenSpot-v2 Desktop | 78.7 | 6.0 | 28.1 | 45.8 | 79.9 | 82.0 | 63.8 | 76.3 |
| GUI | ScreenSpot-v2 Mobile | 81.2 | 7.6 | 42.9 | 60.3 | 86.2 | 87.8 | 69.7 | 81.4 |
| GUI | ScreenSpot-v2 Web | 82.2 | 2.5 | 22.4 | 47.6 | 79.9 | 82.6 | 65.9 | 77.8 |
| Average | - | 69.4 | 57.2 | 52.0 | 59.7 | 64.6 | 69.4 | 63.7 | 70.1 |
*All values in the table are normalized to 0–100. Evaluation is done using vLLM 0.26.0 and each model’s recommended generation parameters when available. Non-reasoning mode is used everywhere, and models are prompted to directly answer without reasoning.
We also evaluated LFM2.5-VL-3B on text-only benchmarks for instruction following and tool use. Instruction following climbs across the board, and tool use improves sharply. On tool use, LFM2.5-VL-3B is on par with Gemma-4-E2B and Qwen3.5-2B.
| Task | Benchmark | LFM2.5-VL-3B (3.1B) | LFM2-VL-3B (3.1B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | InternVL 3.5 2B (2.4B) | InternVL 3.5 4B (4.7B) | Qwen3.5-2B (2.3B) | Qwen3.5-4B (4.7B) |
|---|---|---|---|---|---|---|---|---|---|
| Instruction following | IFEval | 82.3 | 72.9 | 83.0 | 87.9 | 32.4 | 35.4 | 73.6 | 86.2 |
| Instruction following | IFBench | 25.8 | 20.8 | 34.1 | 39.2 | 24.4 | 24.5 | 28.9 | 33.5 |
| Instruction following | Multi-IF | 59.4 | 46.5 | 69.4 | 77.4 | 16.3 | 16.9 | 53.5 | 66.7 |
| Tool use & function calling | ToolSandbox | 59.5 | 26.4 | 56.5 | 61.6 | N/A | N/A | 47.7 | 65.0 |
| Tool use & function calling | BFCL V4 | 32.5 | 20.5 | 33.2 | 40.0 | N/A | N/A | 33.9 | 53.6 |
*InternVL 3.5 models do not support function-calling.
These results demonstrate that LFM2.5-VL-3B is a strong, general-purpose vision-language model. It covers everyday tasks (captioning, visual question answering, document understanding) and is especially good at grounding objects, reading screens and documents, and calling tools.
Inference speed on CPU and GPU
LFM2.5-VL-3B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX.
On-device inference. LFM2.5-VL-3B decodes 228 tokens/s on an M5 Max and 116 tokens/s on a Ryzen AI Max+ 395, and fits in about 3 GB of memory. It even reaches 20 tokens/s on a Galaxy S26 Ultra, so you can run it fully on-device.
GPU inference. LFM2.5-VL-3B keeps latency consistently low and is the fastest on multi-frame inputs.
LFM2.5-VL-3B is also the fastest on output throughput out of all models we tested, reaching about 11K tokens per second at high concurrency. That is roughly 2× the larger 4B-class models and ahead of even the smaller 2B-class models, which adds up to nearly 1B output tokens per day on a single H100.
How to use LFM2.5-VL-3B
Reach for LFM2.5-VL-3B when you need on-device intelligence for high-volume workloads.
Install the latest version of transformers (compatible with transformers>=5.0.0):
%pip install -q torch torchvision accelerate "transformers>=5.10.1"
Then load and run the model:
import torch
from transformers.image_utils import load_image
from transformers import AutoModelForImageTextToText, AutoProcessor
from IPython.display import display
MODEL_ID = "LiquidAI/LFM2.5-VL-3B"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForImageTextToText.from_pretrained(
MODEL_ID,
device_map="auto",
dtype="bfloat16",
)
img_url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/coco_sample.png"
input_image = load_image(img_url)
display(input_image)
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": input_image},
{"type": "text", "text": "Describe this image in two concise sentences."},
],
}
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
do_sample=True,
temperature=0.2,
top_k=50,
repetition_penalty=1.0,
max_new_tokens=256,
)
output = processor.batch_decode(outputs[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0]
print(output)
Two cats are sleeping on a pink couch with two remote controls.
You can find more hands-on examples on how to use LFM2.5-VL3B for multi-image inputs, grounding, OCR, tool calling, and more in our documentation. Check out our release blog for video examples.
LFM2.5-VL-3B demo
Check out this browser demo of LFM2.5-VL-3B powering a vision-capable chat interface. It allows you to take or upload multiple images and let the model interact with them, including grounding, OCR, and tool use.
Get Started
LFM2.5-VL-3B is available on Hugging Face today.
With LFM2.5, we're delivering on our vision of AI that runs anywhere. These models are:
- Download: LFM2.5-VL-3B on Hugging Face.
- Try: run the WebGPU demo in your browser, no setup needed.
- Fine-tune: adapt LFM2.5-VL-3B to your task with our fine-tuning tutorials.
We can't wait to see what you build.
Citation
Please cite this article as:
Liquid AI, "LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge", Liquid AI Blog, Aug 2026.
Or use the BibTeX citation:
@article{liquidAI2026VL3B,
author = {Liquid AI},
title = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-vl-3b},
}
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み