Liquid AI、オンデバイス推論に対応した最小モデル「LFM2.5-230M」を llama.cpp や MLX などと共に公開
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
Liquid AI は、スマートフォンやロボットでのエージェントタスク実行を目的とした同社最小型のオープンウェイトモデル「LFM2.5-230M」を発表し、オンデバイス推論に対応する複数のフレームワークとの互換性を確保した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るSource Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
Liquid AI が、同社史上最小のモデル「LFM2.5-230M」をリリースしました。このモデルは、スマートフォンやロボット、自動化デバイス上でエージェントタスクを実行することを目的に設計されています。
その狙いは明確です。これは汎用的な推論モデルではなく、エッジハードウェア上でのデータ抽出とツール利用に特化したものです。
要点まとめ
- Liquid AI の LFM2.5-230M は同社最小のモデルで、パラメータ数は 2.3 億。オープンウェイトで LFM2 アーキテクチャをベースにしています。
- Galaxy S25 Ultra では秒間 213 トークン、Raspberry Pi 5 では 42 トークンの速度でオンデバイス推論が可能です。
- Qwen3.5-0.8B や Gemma 3 1B よりも大きなモデルに対し、指示の遵守とデータ抽出の精度で上回っています。
- ツール利用やデータ抽出に最適化されており、数学計算、コード生成、クリエイティブライティングには向きません。
- llama.cpp、MLX、vLLM、SGLang、ONNX に対応し、メモリ使用量は 293〜375 MB です。
LFM2.5-230M とは?
LFM2.5-230M は、パラメータ数 2.3 億のテキスト専用モデルです。基盤となるのは LFM2 アーキテクチャで、全 14 レイヤーから構成されています。
そのうち 8 レイヤーはダブルゲート型 LIV コンボリューションブロック、残りの 6 レイヤーはグループ化クエリアテンション(GQA)ブロックです。このハイブリッドなレイアウト設計により、CPU での高速推論を実現しています。
コンテキスト長は 32,768 トークン、語彙サイズは 65,536 です。知識の更新日は 2024 年半ばまでで、英語、中国語、アラビア語、日本語を含む 10 か国語に対応しています。
Liquid AI チームは、オンデバイス推論をサポートする LFM2.5-230M をリリースしました。llama.cpp、MLX、vLLM、SGLang、ONNX に対応しています。
今回は 2 つのチェックポイントが公開されています。LFM2.5-230M-Base はファインチューニング用の事前学習済みモデルで、LFM2.5-230M は汎用的な指令微調整済みのバージョンです。ライセンスは lfm1.0 です。
トレーニングとポストトレーニング
このモデルは 19 トリリオントークンのデータで事前学習されました。これには 32K のコンテキスト拡張フェーズも含まれています。ポストトレーニングでは、以下の 3 つの段階を順に実行します。
まず、より大規模な LFM2.5-350M から知識蒸留(distillation)を行った教師ありファインチューニングです。次に直接選好最適化(DPO)。最後にマルチドメイン強化学習です。これにより、後続の専門化に対する柔軟性を保つことができます。
230M モデルを大規模モデルと競合させられるのは、この蒸留ステップのおかげです。ターゲットとなるタスクにおいて、より大きな LFM2.5-350M の挙動を引き継いでいます。
ベンチマーク結果
Liquid AI チームは、LFM2.5-230M を 10 のベンチマークで評価しました。これらは知識、指令の遵守、データ抽出、ツール使用など多岐にわたります。
指令遵守の結果は以下の通りです。IFEval では LFM2.5-230M が 71.71 を記録し、Qwen3.5-0.8B(59.94)や Gemma 3 1B IT(63.49)を上回りました。IFBench でも 38.40 で両者をリードしています。また、臨床データ抽出テストである CaseReportBench では 22.51 を達成しました。
| モデル | パラメータ数 | IFEval | IFBench | CaseReportBench | BFCLv4 | MMLU-Pro |
|---|---|---|---|---|---|---|
| LFM2.5-230M | 230M | 71.71 | 38.40 | 22.51 | 21.03 | 20.25 |
| LFM2.5-350M | 350M | 76.96 | 40.69 | 32.45 | 21.86 | 20.01 |
| Granite 4.0-H-350M | 350M | 61.27 | 17.22 | 12.44 | 13.28 | 13.14 |
| Qwen3.5-0.8B (Instruct) | 800M | 59.94 | 22.87 | 13.83 | 18.70 | 37.42 |
| Gemma 3 1B IT | 1B | 63.49 | 20.33 | 2.28 | 7.17 | 14.04 |
LFM2.5-230M は指示の遵守やデータ抽出において優れた性能を発揮しますが、広範な知識においては Qwen3.5-0.8B(MMLU-Pro 37.42)に及びません。同モデルの評価は 20.25 に留まり、またエージェントによるツール利用においても一部で弱さを示しています。例えば τ²-Bench Telecom では 5.26 というスコアです。
Liquid AI はこのモデルの限界についても率直に言及しています。高度な推論を要するタスクには推奨していません。具体的には、複雑な数学的計算やコード生成、そして創造的なライティングが該当します。
使用例と具体例
このモデルは主に 2 つの用途で威力を発揮します。
1 つ目は、大規模なデータ抽出パイプラインです。例えば、10 万件の臨床報告書を構造化されたフィールドに変換する処理を想定してください。4 ビット量子化されたモデルであれば、メモリ使用量は 293〜375 MB に収まり、汎用的な CPU でも動作可能です。これにより、ローカル環境でデータを抽出でき、トークンごとの API 利用料も発生しません。
2 つ目は、軽量なオンデバイス型エージェント処理です。音声入力からツール呼び出しを生成するホームオートメーションのハブや、ユーザーのリクエストを適切な機能に振り分けるスマートフォンアシスタントなどが該当します。
初期段階での実装例として、Liquid AI は Unitree G1 という二足歩行ロボットにこのモデルを搭載しました。処理はすべてロボット本体に搭載された NVIDIA Jetson Orin 上で行われます。ここではモデルが「スキル選択層」として機能し、自然言語の指示をツール呼び出しの一連に変換します。これらの呼び出しは、NVIDIA の SONIC フレームワークから提供される低レベルのスキルを実行するものです。
ツール利用の仕組み
LFM2.5 は、以下の 4 つの手順で関数呼び出しをサポートしています。
まず、システムプロンプト内でツールを JSON 形式として定義します。次に、モデルが特殊なトークンの間に Python 風の関数呼び出し文を生成します。その後、その呼び出しを実行して結果を返却し、最後にモデルが通常のテキストで回答を記述します。
デフォルトでは、呼び出しは Python のリスト形式になります。これは <|tool_call_start|> と <|tool_call_end|> トークンの間に配置されます。以下に、ツール JSON を省略したドキュメント通りのパターンを示します。
Copy CodeCopiedUse a different Browser
<|system|>
List of tools: [{"name": "get_candidate_status",
"parameters": {"candidate_id": {"type": "string"}}}]出力結果は以下の通りです。
output = model.generate(
**inputs,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.05,
max_new_tokens=512,
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))Liquid AI はまた、Unsloth と TRL を用いた LoRA による SFT(教師あり微調整)、DPO(直接好意最適化)、GRPO のための微調整レシピも公開しています。これらはすべて Colab ノートブックとして提供されています。
インタラクティブな解説機能については、以下のスクリプトがページのリサイズに対応します。
(function(){
window.addEventListener("message",function(e){
if(e.data&&e.data.type==="lfm-resize"){
var f=document.getElementById("lfm25-frame");
if(f&&e.data.height){f.style.height=e.data.height+"px";}
}
});
})();モデルの重みは Hugging Face で、技術的な詳細やドキュメントも確認できます。ぜひ Twitter をフォローし、15 万人以上の ML 専門家が参加する当社の SubReddit に加入してください。また、ニュースレターへの登録もお忘れなく。
Telegram も利用可能です。今すぐ Telegram チャンネルに参加しましょう。
原文を表示
Liquid AI shipped LFM2.5-230M, it’s the company’s smallest model to date. The release targets a specific job: running agentic tasks on phones, robots, and automation devices. Both the base and instruction-tuned checkpoints are open-weight on Hugging Face.
The pitch is narrow on purpose. This is not a general reasoning model. It is built for data extraction and tool use on edge hardware.
TL;DR
Liquid AI’s LFM2.5-230M is its smallest model yet: 230M params, open-weight, built on LFM2.
Runs on-device at 213 tok/s on a Galaxy S25 Ultra and 42 on a Raspberry Pi 5.
Beats larger models (Qwen3.5-0.8B, Gemma 3 1B) on instruction following and data extraction.
Tuned for tool use and extraction; not for math, code generation, or creative writing.
Day-one support across llama.cpp, MLX, vLLM, SGLang, and ONNX, with a 293–375 MB footprint.
What is LFM2.5-230M?
LFM2.5-230M is a 230-million-parameter, text-only model. It is built on the LFM2 architecture. The model has 14 layers total. Eight are double-gated LIV convolution blocks. The remaining six are grouped-query attention (GQA) blocks. The hybrid layout targets fast CPU inference.
The context length is 32,768 tokens. The vocabulary size is 65,536. The knowledge cutoff is mid-2024. It supports ten languages, including English, Chinese, Arabic, and Japanese.
Liquid AI team ships two checkpoints. LFM2.5-230M-Base is the pre-trained model for fine-tuning. LFM2.5-230M is the general-purpose instruction-tuned version. The license is lfm1.0.
Training and Post-Training
The model was pre-trained on 19 trillion tokens. That total includes a 32K context extension phase. The post-training recipe then runs in three stages.
First comes supervised fine-tuning with distillation from the larger LFM2.5-350M. Second is direct preference optimization (DPO). Third is multi-domain reinforcement learning. This preserves flexibility for downstream specialization.
The distillation step is what keeps a 230M model competitive with larger checkpoints. It inherits behavior from the bigger LFM2.5-350M on targeted tasks.
Benchmark
Liquid AI team evaluated LFM2.5-230M across ten benchmarks. They span knowledge, instruction following, data extraction, and tool use.
The instruction-following results support that. On IFEval, LFM2.5-230M scores 71.71. That beats Qwen3.5-0.8B (59.94) and Gemma 3 1B IT (63.49). On IFBench it scores 38.40, ahead of both. On CaseReportBench, a clinical data-extraction test, it scores 22.51.
ModelParamsIFEvalIFBenchCaseReportBenchBFCLv4MMLU-Pro
LFM2.5-230M230M71.7138.4022.5121.0320.25
LFM2.5-350M350M76.9640.6932.4521.8620.01
Granite 4.0-H-350M350M61.2717.2212.4413.2813.14
Qwen3.5-0.8B (Instruct)800M59.9422.8713.8318.7037.42
Gemma 3 1B IT1B63.4920.332.287.1714.04
LFM2.5-230M leads on instruction following and data extraction. It trails on broad knowledge: MMLU-Pro is 20.25, behind Qwen3.5-0.8B’s 37.42. It is also weak on some agentic tool use. On τ²-Bench Telecom it scores just 5.26.
Liquid AI is direct about the limits. It does not recommend the model for reasoning-heavy workloads. That means advanced math, code generation, and creative writing.
Use Cases With Examples
The model fits two jobs well.
The first is large-scale data extraction pipelines. Picture a pipeline parsing 100,000 clinical reports into structured fields. A 4-bit build with a 293–375 MB memory footprint runs that on commodity CPUs. You extract locally, with no per-token API bill.
The second job is lightweight on-device agentic workloads. Think a home automation hub that turns speech into tool calls. Or a phone assistant that routes a request to the right function.
As an early signal, Liquid AI deployed the model on a Unitree G1 humanoid robot. It ran entirely on the robot’s onboard NVIDIA Jetson Orin. There the model acted as a skill-selection layer. It turned one natural-language instruction into a sequence of tool calls. Those calls invoked low-level skills from NVIDIA’s SONIC framework.
Tool Use: How It Works
LFM2.5 supports function calling in four steps. You define tools as JSON in the system prompt. The model writes a Pythonic function call between special tokens. You execute the call and return the result. The model then writes a plain-text answer.
By default the call is a Python list. It sits between the <|tool_call_start|> and <|tool_call_end|> tokens. Here is the documented pattern, with the tool JSON abbreviated:
Copy CodeCopiedUse a different Browser
<|im_start|>system
List of tools: [{"name": "get_candidate_status",
"parameters": {"candidate_id": {"type": "string"}}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
You can also force JSON-formatted calls through the system prompt.
Running It: A Minimal Example
The model works with Transformers 5.0.0 and up. The recommended generation settings are temperature 0.1, top_k 50, and repetition_penalty 1.05. Note the do_sample=True flag, which is required for those sampling settings to apply.
Copy CodeCopiedUse a different Browser
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "LiquidAI/LFM2.5-230M"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": "What is C. elegans?"}],
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
output = model.generate(
**inputs,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.05,
max_new_tokens=512,
)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Liquid AI also publishes fine-tuning recipes. They cover SFT, DPO, and GRPO with LoRA, via Unsloth and TRL. Each ships as a Colab notebook.
Interactive Explainer
Check out the Model weight on HF, Technical details and Docs. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang, and ONNX Support for On-Device Inference appeared first on MarkTechPost.
関連記事
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み