webAI、ローカル向け論理推論モデル「TwIL-LM」を公開
本文の状態
日本語全文を表示中
詳細モードで約6分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
webAI は、ローカル環境で動作する形式論理推論モデル「TwIL-LM」の1.7Bおよび3B版をリリースし、小規模パラメータでありながら大規模モデルに匹敵する性能を発揮することを発表した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月11日 16:20
AI深層分析
キーポイント
ローカル実行可能な形式論理モデルの公開
webAI は1.7Bと3Bのパラメータを持つ「TwIL-LM」ファミリーをリリースし、英語から一階述語論理への翻訳や推論検証をローカルハードウェアで実行可能にした。
高性能な小規模モデルの技術的アプローチ
3BモデルはSmolLM3-3Bを基にLoRA微調整とチェックポイント融合、MGPO手法を適用して構築され、1.7Bモデルも同様のアーキテクチャを採用している。
大規模モデルとの性能比較における優位性
webAIの発表によると、TwIL-LM3はgpt-oss-120bに対し形式推論の5つの課題のうち4つで上回り、生成トークン数と処理速度でも大幅な効率化を示した。
非商用ライセンスと適用分野
本モデルは現在非商用のみ利用可能であり、コンプライアンス、RegTech、医療、法務などのデータが端末外に出せない環境での活用を想定している。
TwIL-LM3 の両軸での性能向上
TwIL-LM3 はドメイン内スコアが相対的に26%向上し、保持されたコア平均でも0.022の改善を達成した。これはプロジェクト内で両方のトラックで改善を示す唯一のアプローチである。
重要な引用
webAI has released TwIL-LM, a two-model family of formal-logic reasoners at 1.7B and 3B parameters.
TwIL-LM3 produces the shortest generations of any arm... and consequently the most answers per second at 32.9 against the 120B's 4.2.
The model card calls it the only arm in the project that gains on both tracks.
WiSE-FT at λ = 0.25 is why in-domain gains do not collapse held-out performance.
編集コメントを表示
編集コメント
小規模モデルが形式論理という高度な推論タスクで大規模モデルと競合できる点は、エッジAIやローカルAIの実用化において極めて重要な進展である。特にデータ主権が重視される分野での導入ケースが増えることが予想される。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
webAI は、17 億パラメータと 30 億パラメータの 2 つのモデルからなる形式論理推論モデル「TwIL-LM」をリリースしました。30 億パラメータ版(TwIL-LM3)は SmolLM3-3B を統合して微調整したもので、17 億パラメータ版は SmolLM2-1.7B-Instruct に対する PEFT LoRA アダプターです。両モデルの目的は「自動形式化」、つまり英語を一階述語論理(FOL)に変換し、結論が前提から導かれるか検証することです。
どちらもローカル環境で動作します。17 億パラメータ版は量子化済みで 1.06 GB、30 億パラメータ版は Q4_K_M GGUF 形式で 1.78 GiB です。webAI の発表では、5 つの論理推論タスクのうち 4 つで「gpt-oss-120b」を上回る性能を達成したと強調されています。
導入可能でしょうか?
部分的には可能です。ただし現時点では非商用利用に限定されます。
両モデルは「webAI Non-Commercial License ver. 1.0」の下で提供されています。収益化を目的とした展開を行う場合は、別途 webAI と契約を結ぶ必要があります。
企業規模:規模は問いません。30 億パラメータ版の Q4_K_M GGUF は 1.78 GiB で、CPU または VRAM 4 GB で動作します。17 億パラメータ版の Q4_K_M は 1.06 GB です。
適用業界:コンプライアンスと RegTech(規制技術)、金融サービス、ヘルスケア・製薬、法務・契約管理、形式手法研究などです。webAI は、データをデバイス外に出せない環境でのローカル実行を推奨しています。
活用事例:一階述語論理(FOL)への翻訳、前提セットに対する帰結分類、自然言語から構造化クエリへの変換、Lean による形式化のドラフト作成と批判的検討、大規模モデルの出力を検証する検証レイヤーなどです。
TwIL-LM3 はどのように作られたのか?
ベースモデルの上には、4 つの段階が積み重なっています。まず、合成された形式論理コーパスを用いた LoRA による教師あり微調整です。次に、パラメータ空間上で中間 SFT チェックポイントを平均化するチェックポイント融合。さらに、事前学習済みベースモデルへ λ = 0.25 の重みで補間する WiSE-FT です。最後に、プログラム検証器に対して実行されるエントロピー加重 GRPO 段階の MGPO が続きます。公開されているチェックポイントはステップ 2071 です。
このλは非常に重要です。微調整されたデルタのうち、わずか 4 分の 1 しか保持されません。補間をスキップした兄弟モデルの方がドメイン内では高いスコア(マクロゲートで 0.515)を示しましたが、その代わり、保持された能力が約 12 ポイント低下しました。webAI はこのモデルは公開していません。
性能について
webAI の発表によると、ルール推論で 96.4、意味解析で 87.6、Lean 形式化で 64.6、正確なフォーマットでの回答で 52.0、含意ラベリングで 68.7 を記録しています。
この評価は 2 つのトラックに分かれています。トラック A(ドメイン内の形式論理)では、TwIL-LM3 は 6 ランの平均で 0.4488、トレーニングパイプラインがゲートとする指標であるマクロゲートで 0.4218 を達成しました。これは、LFM2.5-8B-A1B を含むすべてのモデルを上回る結果です。パラメータ数は約 3 分の 1 ですが、6 つの目的レーンすべてで 0.4218 というスコアを叩き出し、対照的な 0.3757 を上回っています。ただし、最も大きな 2 つのモデルには勝っていません。
Qwen3-8B はマクロゲートで 0.5336 を記録し、TwIL-LM3 の 0.4218 を上回っていますが、その多くは緩い一致による加点です。厳格な評価(strict-7)では両者とも 0.2093 と 0.1971 に落ち着きます。また、gpt-oss-120b は 6 ランの平均で 0.5192 を記録し、TwIL-LM3 の 0.4488 を上回っています。
モデルカードにおいて明確に示されているのは、効率性です。TwIL-LM3 はあらゆるアームの中で最も短い生成結果(Track B で 482 トークン)を出力し、その結果、1 秒あたりの回答数は 32.9 と、120B モデルの 4.2 を大きく上回ります。
imagehttps://www.webai.com/blog/webai-releases-twil-lm-a-family-of-formal-logic-models-that-outreason-a-120b-model-and-run-on-an-iphone
ホールドアウト転移性能
TwIL-LM3 は、ドメイン内での性能が相対的に 26% 向上し、マクロゲート値は 0.336 から 0.422 に改善しました。同時に、ホールドアウトコアの平均スコアも 0.022 上昇しています。モデルカードでは、このプロジェクトにおいて両方のトラックで性能が向上した唯一のアームであると紹介されています。
LogicBench のスコアは 0.6467 から 0.7167 に向上しました。一方、GSM8K はわずかに低下し 0.8833 から 0.8733 へ、IFEval も 0.6767 から 0.6433 へと後退しています。
1.7B モデルのトレードオフ
1.7B モデルは異なるトレードオフを示します。そのマクロプライマリースコアは、未適応ベースラインの 0.185 を大きく上回る 0.361 です。ただし、分布外(Out-of-distribution)の結果は混合しています。
LogicBench BQA は 0.563 から 0.590 に改善しましたが、GSM8K は 0.413 から 0.380 に低下し、ARC-C の chain-of-thought も 0.587 から 0.463 へと下がっています。
キーポイント
- TwIL-LM3(3B)と TwIL-LM(1.7B)は形式論理を目的としており、どちらも非商用ライセンスの下で提供されます。
- TwIL-LM3 は、6 つのトラックの平均において gpt-oss-120b に劣ります(0.4488 対 0.5192)。
- その真の強みは効率性です。482 トークンの生成から 1 秒あたり 32.9 の回答を出力します。
- WiSE-FT を λ = 0.25 に設定することで、ドメイン内の性能向上がホールドアウト性能の低下を引き起こさないようになっています。
モデルの重みと技術詳細は、こちらからご確認ください。また、Twitter でフォローしていただくことも可能です。15 万人以上の ML 関連ユーザーが参加する SubReddit にぜひご入会ください。さらに、ニュースレターへの登録も忘れずに。
Telegram をご利用の方へ:今なら Telegram でもコミュニティに参加できるようになりました。
この記事は MarkTechPost にて公開された「webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware」です。
原文を表示
webAI has released TwIL-LM, a two-model family of formal-logic reasoners at 1.7B and 3B parameters. The 3B member, TwIL-LM3, is a merged fine-tune of SmolLM3-3B; the 1.7B member is a PEFT LoRA adapter for SmolLM2-1.7B-Instruct. Both target autoformalization: translating English into first-order logic and checking whether a conclusion follows from its premises. Both run locally, with a 1.06 GB quantized build for the 1.7B and a 1.78 GiB Q4_K_M GGUF for the 3B. webAI’s announcement frames the release around beating gpt-oss-120b on four of five formal-reasoning lanes.
Is it deployable?
Partially. Non-commercial use only, as of now.
Both checkpoints ship under the webAI Non-Commercial License ver. 1.0. Revenue-generating deployment requires a separate agreement with webAI.
Company level: any size. The 3B Q4_K_M GGUF is 1.78 GiB and runs on CPU or 4 GB of VRAM. The 1.7B Q4_K_M is 1.06 GB.
Industries: compliance and RegTech, financial services, healthcare and pharma, legal and contract operations, formal-methods research. webAI positions local execution for environments where data cannot leave the device.
Applications: first-order logic (FOL) translation, entailment classification over premise sets, natural language to structured query, Lean formalization drafting and critique, and a verifier layer that checks a larger model’s output.
How TwIL-LM3 was built?
Four stages sit on top of the base model. LoRA supervised fine-tuning on a synthetic formal-logic corpus. Checkpoint fusion, averaging intermediate SFT checkpoints in parameter space. WiSE-FT interpolation back toward the pretrained base at λ = 0.25. Then MGPO, an entropy-weighted GRPO stage run against a programmatic verifier. The published checkpoint is step 2071.
That λ is load-bearing: only a quarter of the fine-tuned delta is retained. A sibling arm that skipped the interpolation scored higher in-domain, at macro gate 0.515, but gave back roughly twelve points of held-out capability. webAI did not publish that arm.
Performance
webAI's announcement lists 96.4 on rule induction, 87.6 on semantic parsing, 64.6 on Lean formalization, 52.0 on exact-format answering, and 68.7 on entailment labeling.
It reports two tracks. On Track A, in-domain formal logic, TwIL-LM3 scores 0.4488 on the six-lane average and 0.4218 on the macro gate, the metric the training pipeline gates on. It leads every arm up to and including LFM2.5-8B-A1B on all six objective lanes, at 0.4218 against 0.3757 with a third of the parameters. It does not lead the two largest arms. Qwen3-8B takes the gate 0.5336 to 0.4218, but most of that is loose-match credit; under strict-7 the two sit at 0.2093 and 0.1971. gpt-oss-120b takes the six-lane average 0.5192 to 0.4488.
Efficiency is where the model card is unambiguous. TwIL-LM3 produces the shortest generations of any arm, 482 tokens on Track B, and consequently the most answers per second at 32.9 against the 120B's 4.2.
imagehttps://www.webai.com/blog/webai-releases-twil-lm-a-family-of-formal-logic-models-that-outreason-a-120b-model-and-run-on-an-iphone
Held-out transfer
TwIL-LM3 improves in-domain by +26% relative, macro gate 0.336 to 0.422, while also gaining +0.022 on the held-out core average. The model card calls it the only arm in the project that gains on both tracks. LogicBench moves to 0.7167 from 0.6467. GSM8K slips slightly to 0.8733 from 0.8833, and IFEval regresses to 0.6433 from 0.6767.
The 1.7B is a different trade. Its macro-primary score is 0.361 against 0.185 for the unadapted base. Out-of-distribution results are mixed: LogicBench BQA improves to 0.590 from 0.563, while GSM8K falls to 0.380 from 0.413 and ARC-C chain-of-thought falls to 0.463 from 0.587.
Key Takeaways
TwIL-LM3 (3B) and TwIL-LM (1.7B) target formal logic, both under a non-commercial license.
Shipping TwIL-LM3 trails gpt-oss-120b on the six-lane average, 0.4488 to 0.5192.
Its real edge is efficiency: 32.9 answers/sec from 482-token generations.
WiSE-FT at λ = 0.25 is why in-domain gains do not collapse held-out performance.
Check out the Model weights and Technical details. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
The post webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware appeared first on MarkTechPost.
関連記事
News to Guide
ニュースの次に確認する
発表内容を、現在の料金や仕様と照らし合わせられる関連ガイドです。
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み