IMDb感情分析:DistilBERT LoRAとTF-IDF比較・校正・解釈性テスト
本文の状態
日本語全文を表示中
詳細モードで約12分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
MarkTechPost
Stanford NLPのIMDb映画レビューデータセットを用い、TF-IDFベースラインとLoRAによるDistilBERT微調整を比較するエンドツーエンドの感情分析ワークフローを開発し、データ監査や校正、解釈性評価を含む手法を示した。
Continue in AI NEW LAB
このニュースを、実務の判断につなげる
AI NEW LABで、試したことや先に確認したい条件を共有できます。まずはログインなしで読めます。
AI NEW LABで論点を見るAI深層分析を開く2026年8月9日 16:21
AI深層分析
キーポイント
厳格なデータ監査とベースライン構築
クラス順序、レビュー長さの偏り、重複漏洩などのデータを精査し、TF-IDF とロジスティック回帰による強固なベースラインを確立した。
LoRA を用いた DistilBERT の微調整と評価
PEFT 経由で DistilBERT を LoRA で微調整し、精度や ROC-AUC だけでなく、閾値選択や確率較正(ECE)も詳細に評価した。
モデルの解釈可能性と頑健性の検証
単語レベルのオクルージョンサリエンスや文頭・文末の切り捨て影響を調査し、長文コンテキストにおける限界や予測根拠を解明した。
半教師あり学習による性能向上
ラベルなしデータを用いた信頼度ベースの疑似ラベリングを実行し、従来のベースラインと比較してモデル性能を向上させた。
環境設定とライブラリ自動インストール
必要なライブラリが不足している場合、コード内で自動的にインストールする仕組みを実装し、シード値を固定して実験の再現性を確保している。
重要な引用
In this tutorial, we develop an end-to-end sentiment analysis workflow using the Stanford NLP IMDb Large Movie Review Dataset and compare classical machine learning with parameter-efficient transformer fine-tuning.
Beyond headline metrics, we investigate confident errors, performance across review lengths, word-level occlusion saliency, and head-versus-tail truncation to understand how the model reaches its predictions.
print(f"Installing: {', '.join(_missing)} ...")
raw = load_dataset("stanfordnlp/imdb")
編集コメントを表示
編集コメント
本記事は、最新の Transformer 技術を応用しつつも古典的な機械学習手法との比較を丁寧に行っており、実務におけるモデル選定や評価の基準となる貴重なガイドである。特にデータ監査と解釈可能性分析に言及している点は、実運用を意識する開発者にとって極めて有益な内容と言える。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
本チュートリアルでは、Stanford NLP の IMDb 大規模映画レビューデータセットを用いて、感情分析のワークフローをゼロから構築します。古典的な機械学習手法と、パラメータ効率的なトランスフォーマーのファインチューニング(LoRA)を比較検討します。
まず、再現性の高い環境を整え、トレーニング前にデータセットのクラス順序、レビュー長さの偏り、重複によるリーク、前処理に伴うアーティファクトなどを精査します。その上で、TF-IDF とロジスティック回帰を組み合わせた強力なベースラインモデルを構築します。
次に、PEFT を活用して DistilBERT に LoRA でファインチューニングを施し、精度(accuracy)、マクロ F1 スコア、ROC-AUC、混同行列、ROC カーブブなどの指標で評価を行います。さらに、閾値の選択や確率のキャリブレーションについて、Expected Calibration Error と信頼性分析を通じて詳細に検証します。
単なる主要な数値指標を超えて、モデルが予測に至るプロセスと、文脈長(long-context)の制限が性能にどう影響するかを解明するため、自信過剰な誤り(confident errors)、レビュー長さごとの性能差、単語レベルのオクルージョン・サリエンス分析、および先頭・末尾の切り捨て(truncation)の影響についても調査します。
最後に、ラベルなしの IMDb データセットを用いて信頼度に基づく疑似ラベリング(pseudo-labeling)を行い、得られた半教師ありモデルをベースラインと比較。学習済みのトランスフォーマーをマージして、再利用可能な感情分析推論モデルとして保存します。
import importlib.util, subprocess, sys, os, time, random, warnings, inspect, hashlib
warnings.filterwarnings("ignore")
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["WANDB_DISABLED"] = "true"
_REQUIRED = {
"transformers": "transformers",
"datasets": "datasets",
"peft": "peft",
"accelerate": "accelerate",
"sklearn": "scikit-learn",
}
_missing = [pkg for mod, pkg in _REQUIRED.items() if importlib.util.find_spec(mod) is None]
if _missing:
print(f"Installing: {', '.join(_missing)} ...")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", *_missing], check=True)
print("Done. (If imports fail below, restart the runtime and re-run.)\n")
import numpy as np
import pandas as pd
import torch
import matplotlib.pyplot as plt
from datasets import load_dataset
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import (accuracy_score, f1_score, roc_auc_score,
classification_report, confusion_matrix, roc_curve)
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer, DataCollatorWithPadding,
EarlyStoppingCallback, set_seed)
from peft import LoraConfig, get_peft_model, TaskType
def _disable_torchao_probe():
patched = []
try:
import peft.import_utils as _piu
_piu.is_torchao_available = lambda: False
patched.append("peft.import_utils")
except Exception:
pass
for _name, _mod in list(sys.modules.items()):
if _name.startswith("peft") and hasattr(_mod, "is_torchao_available"):
_mod.is_torchao_available = lambda: False
patched.append(_name)
return patched
try:
import torchao as _tao
_v = getattr(_tao, "__version__", "?")
if tuple(int(x) for x in _v.split(".")[:2]) disabling PEFT's torchao probe: "
f"{', '.join(_disable_torchao_probe())}")
except Exception:
_disable_torchao_probe()
SEED = 42
MODEL_NAME = "distilbert-base-uncased"
MAX_LEN = 256
N_TRAIN = 5000
N_EVAL = 2000
N_UNSUP = 3000
EPOCHS = 2
BATCH = 16
LR = 3e-4
FULL_RUN = False
if FULL_RUN:
N_TRAIN, N_EVAL, EPOCHS = 25000, 25000, 3
set_seed(SEED); random.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
print("=" * 79)
print(f"device={DEVICE} | torch={torch.__version__} | "
f"gpu={torch.cuda.get_device_name(0) if DEVICE=='cuda' else 'n/a'}")
print("=" * 79)
t0 = time.time()
raw = load_dataset("stanfordnlp/imdb")
print(raw, f"\nloaded in {time.time()-t0:.1f}s\n")
print("--- example (truncated) ---")
print("label:", raw["train"][0]["label"], "|", raw["train"][0]["text"][:300], "...\n")
first_labels = np.array(raw["train"]["label"][:5])
last_labels = np.array(raw["train"]["label"][-5:])
print(f"TRAP #1 - split ordering: first 5 labels {first_labels}, "
f"last 5 labels {last_labels} -> ALWAYS shuffle before subsampling.")
train_full = raw["train"].shuffle(seed=SEED)
test_full = raw["test"].shuffle(seed=SEED)
train_ds = train_full.select(range(min(N_TRAIN, len(train_full))))
eval_ds = test_full.select(range(min(N_EVAL, len(test_full))))
print(f" after shuffle+subsample: train balance = "
f"{np.bincount(train_ds['label'])}, eval balance = {np.bincount(eval_ds['label'])}")
lens = np.array([len(t.split()) for t in train_full["text"]])
q = np.percentile(lens, [50, 75, 90, 95, 99])
print(f"\nTRAP #2 - length (words): median={q[0]:.0f} p75={q[1]:.0f} p90={q[2]:.0f} "
f"p95={q[3]:.0f} p99={q[4]:.0f} max={lens.max()}")
print(f" ~{(lens > MAX_LEN*0.75).mean()*100:.1f}% of reviews exceed MAX_LEN={MAX_LEN} "
f"tokens (rough words->tokens factor 1.3). Section 9 measures what that costs.")
h_tr = {hashlib.md5(t.encode()).hexdigest() for t in raw["train"]["text"]}
h_te = {hashlib.md5(t.encode()).hexdigest() for t in raw["test"]["text"]}
print(f"\nTRAP #3 - leakage: {len(h_tr & h_te)} exact duplicate reviews across "
f"train/test; {len(raw['train'])-len(h_tr)} dupes inside train itself.")
def clean(t):
return t.replace(
", " ").replace(
", " ").strip()
plt.figure(figsize=(11, 3.2))
plt.subplot(1, 2, 1)
plt.hist(np.clip(lens, 0, 1000), bins=60)
plt.axvline(MAX_LEN, ls="--", color="k", label=f"MAX_LEN={MAX_LEN}")
plt.title("Review length (words, clipped at 1000)"); plt.legend()
plt.subplot(1, 2, 2)
plt.bar(["neg", "pos"], np.bincount(raw["train"]["label"]))
plt.title("Train class balance (perfectly balanced)")
plt.tight_layout(); plt.show()
Colab の環境を設定し、必要なライブラリをインストール。PEFT と torchao の互換性パッチを適用して実験の再現性を確保するため、乱数シードを固定します。Stanford IMDb データセットを読み込み、訓練用とテスト用のデータをシャッフル・サンプリング。クラスバランスやレビュー長さの分布、重複データによるリーク、HTML 文字列の混在などを確認します。モデル構築前にデータ構造を理解できるよう、レビュー長さとラベル頻度を可視化しておきましょう。
Copy CodeCopiedUse a different Browser
print("\n" + "=" * 79 + "\n3. TF-IDF BASELINE\n" + "=" * 79)
Xtr = [clean(t) for t in train_ds["text"]]; ytr = np.array(train_ds["label"])
Xte = [clean(t) for t in eval_ds["text"]]; yte = np.array(eval_ds["label"])
t0 = time.time()
tfidf_clf = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_features=300_000,
sublinear_tf=True, strip_accents="unicode"),
LogisticRegression(C=8.0, max_iter=2000, n_jobs=-1),
)
tfidf_clf.fit(Xtr, ytr)
p_tfidf = tfidf_clf.predict_proba(Xte)[:, 1]
acc_tfidf = accuracy_score(yte, p_tfidf > 0.5)
auc_tfidf = roc_auc_score(yte, p_tfidf)
print(f"trained in {time.time()-t0:.1f}s -> acc={acc_tfidf:.4f} auc={auc_tfidf:.4f}")
vec, lr = tfidf_clf.steps[0][1], tfidf_clf.steps[1][1]
feats, coefs = np.array(vec.get_feature_names_out()), lr.coef_[0]
order = np.argsort(coefs)
print("\nmost NEGATIVE n-grams:", ", ".join(feats[order[:12]]))
print("most POSITIVE n-grams:", ", ".join(feats[order[-12:]][::-1]))
print("\n" + "=" * 79 + "\n4. LoRA FINE-TUNING\n" + "=" * 79)
tok = AutoTokenizer.from_pretrained(MODEL_NAME)
def tokenize(batch):
return tok([clean(t) for t in batch["text"]], truncation=True, max_length=MAX_LEN)
tr_tok = (train_ds.map(tokenize, batched=True, remove_columns=["text"])
.rename_column("label", "labels"))
ev_tok = (eval_ds.map(tokenize, batched=True, remove_columns=["text"])
.rename_column("label", "labels"))
base = AutoModelForSequenceClassification.from_pretrained(
MODEL_NAME, num_labels=2,
id2label={0: "NEGATIVE", 1: "POSITIVE"},
label2id={"NEGATIVE": 0, "POSITIVE": 1},
)
lora_cfg = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_lin", "v_lin"],
modules_to_save=["pre_classifier", "classifier"],
)
try:
model = get_peft_model(base, lora_cfg)
except ImportError as e:
_disable_torchao_probe()
print(f"[compat] retrying after backend probe failure: {e}")
model = get_peft_model(base, lora_cfg)
model.print_trainable_parameters()
def compute_metrics(eval_pred):
logits, labels = eval_pred
probs = torch.softmax(torch.tensor(logits), dim=-1).numpy()[:, 1]
preds = (probs > 0.5).astype(int)
return {"accuracy": accuracy_score(labels, preds),
"f1_macro": f1_score(labels, preds, average="macro"),
"roc_auc": roc_auc_score(labels, probs)}
_ta = inspect.signature(TrainingArguments.__init__).parameters
_eval_key = "eval_strategy" if "eval_strategy" in _ta else "evaluation_strategy"
ta_kwargs = dict(
output_dir="./imdb_lora", learning_rate=LR,
per_device_train_batch_size=BATCH, per_device_eval_batch_size=BATCH * 2,
num_train_epochs=EPOCHS, weight_decay=0.01, warmup_ratio=0.06,
logging_steps=50, save_strategy="epoch", save_total_limit=1,
load_best_model_at_end=True, metric_for_best_model="accuracy",
fp16=(DEVICE == "cuda"), report_to="none", seed=SEED,
)
ta_kwargs[_eval_key] = "epoch"
Trainer の初期化パラメータを検証し、processing_class が存在する場合はそのキーを、そうでない場合は tokenizer を使用して設定します。その後、モデルとトレーニング引数(TrainingArguments)を指定し、学習用および評価用のトークナイズ済みデータセットを割り当てます。パディング処理には DataCollatorWithPadding を利用し、計算メトリクスとして compute_metrics を登録します。また、早期停止(patience=2)を有効にする EarlyStoppingCallback をコールバックリストに追加し、トークナイザーを動的に設定して Trainer インスタンスを作成します。
トレーニング開始時刻を記録し、trainer.train() を実行して学習を開始します。終了後には、所要時間を分単位で表示するメッセージを出力します。
解釈可能な参照点として、強力な TF-IDF とロジスティック回帰のベースラインモデルを訓練し、最も影響力のある正負の n-gram を分析します。次に、IMDb のレビューをトークナイズし、DistilBERT に LoRA アダプターを設定して、バックボーン(骨格)をほぼ凍結したままパラメータの一部のみを更新する構成にします。
Hugging Face の Trainer を用いて、動的パディング、早期停止、混合精度計算、複数の評価指標を活用し、トランスフォーマーモデルの効率的なファインチューニングを行います。
print("\n" + "=" * 79 + "\n5. EVALUATION\n" + "=" * 79)
pred_out = trainer.predict(ev_tok)
p_lora = torch.softmax(torch.tensor(pred_out.predictions), dim=-1).numpy()[:, 1]
y_true = np.array(pred_out.label_ids)
yhat = (p_lora > 0.5).astype(int)
print(classification_report(y_true, yhat, target_names=["neg", "pos"], digits=4))
cm = confusion_matrix(y_true, yhat)
fig, ax = plt.subplots(1, 2, figsize=(11, 4))
ax[0].imshow(cm, cmap="Blues")
for i in range(2):
for j in range(2):
ax[0].text(j, i, cm[i, j], ha="center", va="center", fontsize=14)
ax[0].set_xticks([0, 1], ["pred neg", "pred pos"])
ax[0].set_yticks([0, 1], ["true neg", "true pos"]); ax[0].set_title("Confusion matrix")
for name, p in [("TF-IDF", p_tfidf), ("DistilBERT+LoRA", p_lora)]:
fpr, tpr, _ = roc_curve(y_true, p)
ax[1].plot(fpr, tpr, label=f"{name} (AUC={roc_auc_score(y_true, p):.4f})")
ax[1].plot([0, 1], [0, 1], "k--", lw=0.8)
ax[1].set_xlabel("FPR"); ax[1].set_ylabel("TPR"); ax[1].set_title("ROC"); ax[1].legend()
plt.tight_layout(); plt.show()
print("\n" + "=" * 79 + "\n6. THRESHOLD & CALIBRATION\n" + "=" * 79)
ths = np.linspace(0.05, 0.95, 91)
accs = [(y_true == (p_lora > t)).mean() for t in ths]
best_t = ths[int(np.argmax(accs))]
print(f"acc@0.50 = {accs[45]:.4f} | best threshold = {best_t:.2f} -> acc = {max(accs):.4f}")
def expected_calibration_error(probs, labels, n_bins=10):
"""ECE: |confidence - accuracy| averaged over confidence bins."""
conf = np.maximum(probs, 1 - probs)
正解判定は、確率が 0.5 を超える場合に 1 とし、ラベルと比較することで行います。
バインディング(区間)を 0 から 1 の範囲で n_bins+1 個作成し、期待不一致誤差 (ECE) の計算を開始します。各バインダの下限と上限に対して、信頼度がその範囲内にあるサンプルを抽出し、平均正解率と平均信頼度を記録していきます。
次に、モデルが最も自信を持って間違えたケース(3 つ)を表示します。これは、予測結果がラベルと一致せず、かつモデルの自信度が高い事例です。各事例について、真値(pos/neg)、予測値、信頼度の数値、および単語数を出力し、テキスト内容も表示します。
さらに、レビューの長さによる精度の変化を分析します。サンプルを「短い」「中程度」「長い」「非常に長い」の 4 つのバケットに分類し、各バケットごとの平均正解率とサンプル数を算出します。この結果から、長文レビューでは切り捨て処理が精度に悪影響を与えることが確認できます。
=== 8. オクルージョン・サリエンス分析 ===
推論用モデルとして、LoRA パラメータを統合して元の重みに戻したモデルを用意し、指定されたデバイス上で評価モードに設定します。
確率予測を行う関数では、入力テキストをバッチ処理でトークン化し、最大長まで切り捨てながらパディングを行います。その後、モデルのロジットに対してソフトマックスを適用し、正クラス(pos)の確率を取得して返却します。
オクルージョン分析の関数では、テキスト内の単語を順次隠蔽(マスク)しながら、モデルの予測がどのように変化するかを検証します。
words = clean(text).split()[:max_words]
base = predict_proba([" ".join(words)])[0]
variants = [" ".join(words[:i] + words[i + 1:]) for i in range(len(words))]
dropped = predict_proba(variants)
return words, base - dropped, base
sample = err[err.correct].nlargest(1, "confidence").iloc[0]
words, contrib, base_p = occlusion(sample.text)
print(f"P(positive) for the full excerpt = {base_p:.3f} "
f"(true label = {'pos' if sample.y else 'neg'})\n")
top = np.argsort(np.abs(contrib))[-15:]
plt.figure(figsize=(7, 5))
plt.barh(range(len(top)), contrib[top],
color=["tab:green" if contrib[i] > 0 else "tab:red" for i in top])
plt.yticks(range(len(top)), [words[i] for i in top])
plt.xlabel("Δ P(positive) when the word is removed")
plt.title("Occlusion saliency — green pushes POSITIVE, red pushes NEGATIVE")
plt.tight_layout(); plt.show()
print("\n" + "=" * 79 + "\n9. HEAD vs TAIL TRUNCATION\n" + "=" * 79)
probe = err.nlargest(600, "n_words")
W = 180
head_txt = [" ".join(clean(t).split()[:W]) for t in probe.text]
tail_txt = [" ".join(clean(t).split()[-W:]) for t in probe.text]
yp = probe.y.values
acc_head = ((predict_proba(head_txt) > 0.5).astype(int) == yp).mean()
acc_tail = ((predict_proba(tail_txt) > 0.5).astype(int) == yp).mean()
print(f"on the {len(probe)} longest reviews, using only {W} words:")
print(f" first {W} words -> acc {acc_head:.4f}")
print(f" last {W} words -> acc {acc_tail:.4f}")
実用的な知見:末尾の情報が重要である場合、モデルに「先頭+末尾」を入力するか、MAX_LEN を引き上げるべきです。左側から盲目的に切り捨てるのは避けてください。
ここでは、モデルが最も自信を持って誤ったと判断した予測事例を調査し、レビュー長さに基づいて分類することで、切り捨てに関連する失敗パターンや困難な例を特定します。LoRA アダプターを基盤モデルに統合した後、「1 語除外オクルージョン」手法を適用して、各単語が個別の予測を正または負の感情へと押しやる影響力を推定します。さらに、長いレビューの先頭部分と末尾部分をそれぞれ比較し、最も強い感情情報がどこに含まれているかを明らかにします。
print("\n" + "=" * 79 + "\n10. PSEUDO-LABELLING\n" + "=" * 79)
unsup = raw["unsupervised"].shuffle(seed=SEED).select(range(N_UNSUP))
p_uns = predict_proba(unsup["text"])
keep = (p_uns > 0.95) | (p_uns < 0.05)
pl_labels = np.where(keep, p_uns.argmax(axis=1), -1).astype(int)
print(f"kept {keep.sum()}/{N_UNSUP} pseudo-labels at conf>0.95 "
f"(balance: {np.bincount(pl_labels)})")
aug = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_features=300_000,
sublinear_tf=True, strip_accents="unicode"),
LogisticRegression(C=8.0, max_iter=2000, n_jobs=-1),
).fit(Xtr + pl_texts, np.concatenate([ytr, pl_labels]))
acc_aug = accuracy_score(yte, aug.predict(Xte))
print(f"TF-IDF baseline : {acc_tfidf:.4f}")
print(f"TF-IDF + pseudo-labels: {acc_aug:.4f} (Δ {acc_aug-acc_tfidf:+.4f})")
print("Caveat: gains are bounded by the teacher. Self-training also amplifies "
"the teacher's biases — always validate on clean, held-out data.")
print("\n" + "=" * 79 + "\n11. SAVE & INFER\n" + "=" * 79)
SAVE_DIR = "./imdb-distilbert-lora-merged"
infer_model.save_pretrained(SAVE_DIR); tok.save_pretrained(SAVE_DIR)
print(f"saved merged model to {SAVE_DIR}/ (load with "
f"AutoModelForSequenceClassification.from_pretrained('{SAVE_DIR}'))")
demos = [
"A masterclass in tension. The final act left the whole theatre silent.",
"Two hours I will never get back. Wooden acting, incoherent plot.",
"It's not the disaster the trailer promised, but it never really lands either.",
]
for d, p in zip(demos, predict_proba(demos)):
print(f" P(pos)={p:.3f} -> {'POSITIVE' if p > 0.5 else 'NEGATIVE'} | {d}")
print("\n" + "=" * 79)
print(f"SUMMARY (n_train={N_TRAIN}, n_eval={N_EVAL}, max_len={MAX_LEN})")
print("=" * 79)
print(pd.DataFrame([
{"model": "TF-IDF + LogReg", "accuracy": acc_tfidf, "roc_auc": auc_tfidf},
{"model": "TF-IDF + pseudo-labels", "accuracy": acc_aug, "roc_auc": float("nan")},
{"model": "DistilBERT + LoRA", "accuracy": accuracy_score(y_true, yhat),
"roc_auc": roc_auc_score(y_true, p_lora)},
]).to_string(index=False))
print("""
NEXT EXPERIMENTS
- Set FULL_RUN = True for the real 25k/25k benchmark (~40 min on a T4).
- Swap MODEL_NAME to 'roberta-base' (target_modules=['query','value']) or
'answerdotai/ModernBERT-base' for an 8k context window — no truncation.
- Head+tail truncation: first 128 + last 128 tokens, motivated by section 9.
- Ablate LoRA rank r in {4, 8, 16, 64} and plot accuracy vs trainable params.
- Replace the pseudo-label teacher with an ensemble and iterate self-training.
- Push to the Hub: huggingface_hub.login() then infer_model.push_to_hub(...).
""")
微調整したトランスフォーマーモデルを用いて、IMDb のラベルなしデータセットから高信頼度の疑似ラベルを生成し、これを TF-IDF の訓練コーパスに追加します。
原文を表示
In this tutorial, we develop an end-to-end sentiment analysis workflow using the Stanford NLP IMDb Large Movie Review Dataset and compare classical machine learning with parameter-efficient transformer fine-tuning. We begin by establishing a reproducible environment and auditing the dataset for class ordering, review-length skew, duplicate leakage, and preprocessing artifacts before training a strong TF-IDF and Logistic Regression baseline. We then fine-tune DistilBERT with LoRA through PEFT, evaluate it using accuracy, macro-F1, ROC-AUC, confusion matrices, and ROC curves, and examine threshold selection and probability calibration through Expected Calibration Error and reliability analysis. Beyond headline metrics, we investigate confident errors, performance across review lengths, word-level occlusion saliency, and head-versus-tail truncation to understand how the model reaches its predictions and where long-context limitations affect performance. Finally, we use the unlabeled IMDb split for confidence-based pseudo-labeling, compare the resulting semi-supervised model against our baseline, and save the merged transformer for reusable sentiment inference.
Copy CodeCopiedUse a different Browser
import importlib.util, subprocess, sys, os, time, random, warnings, inspect, hashlib
warnings.filterwarnings("ignore")
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["WANDB_DISABLED"] = "true"
_REQUIRED = {
"transformers": "transformers",
"datasets": "datasets",
"peft": "peft",
"accelerate": "accelerate",
"sklearn": "scikit-learn",
}
_missing = [pkg for mod, pkg in _REQUIRED.items() if importlib.util.find_spec(mod) is None]
if _missing:
print(f"Installing: {', '.join(_missing)} ...")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", *_missing], check=True)
print("Done. (If imports fail below, restart the runtime and re-run.)\n")
import numpy as np
import pandas as pd
import torch
import matplotlib.pyplot as plt
from datasets import load_dataset
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.metrics import (accuracy_score, f1_score, roc_auc_score,
classification_report, confusion_matrix, roc_curve)
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
TrainingArguments, Trainer, DataCollatorWithPadding,
EarlyStoppingCallback, set_seed)
from peft import LoraConfig, get_peft_model, TaskType
def _disable_torchao_probe():
patched = []
try:
import peft.import_utils as _piu
_piu.is_torchao_available = lambda: False
patched.append("peft.import_utils")
except Exception:
pass
for _name, _mod in list(sys.modules.items()):
if _name.startswith("peft") and hasattr(_mod, "is_torchao_available"):
_mod.is_torchao_available = lambda: False
patched.append(_name)
return patched
try:
import torchao as _tao
_v = getattr(_tao, "__version__", "?")
if tuple(int(x) for x in _v.split(".")[:2]) disabling PEFT's torchao probe: "
f"{', '.join(_disable_torchao_probe())}")
except Exception:
_disable_torchao_probe()
SEED = 42
MODEL_NAME = "distilbert-base-uncased"
MAX_LEN = 256
N_TRAIN = 5000
N_EVAL = 2000
N_UNSUP = 3000
EPOCHS = 2
BATCH = 16
LR = 3e-4
FULL_RUN = False
if FULL_RUN:
N_TRAIN, N_EVAL, EPOCHS = 25000, 25000, 3
set_seed(SEED); random.seed(SEED); np.random.seed(SEED); torch.manual_seed(SEED)
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
print("=" * 79)
print(f"device={DEVICE} | torch={torch.__version__} | "
f"gpu={torch.cuda.get_device_name(0) if DEVICE=='cuda' else 'n/a'}")
print("=" * 79)
t0 = time.time()
raw = load_dataset("stanfordnlp/imdb")
print(raw, f"\nloaded in {time.time()-t0:.1f}s\n")
print("--- example (truncated) ---")
print("label:", raw["train"][0]["label"], "|", raw["train"][0]["text"][:300], "...\n")
first_labels = np.array(raw["train"]["label"][:5])
last_labels = np.array(raw["train"]["label"][-5:])
print(f"TRAP #1 - split ordering: first 5 labels {first_labels}, "
f"last 5 labels {last_labels} -> ALWAYS shuffle before subsampling.")
train_full = raw["train"].shuffle(seed=SEED)
test_full = raw["test"].shuffle(seed=SEED)
train_ds = train_full.select(range(min(N_TRAIN, len(train_full))))
eval_ds = test_full.select(range(min(N_EVAL, len(test_full))))
print(f" after shuffle+subsample: train balance = "
f"{np.bincount(train_ds['label'])}, eval balance = {np.bincount(eval_ds['label'])}")
lens = np.array([len(t.split()) for t in train_full["text"]])
q = np.percentile(lens, [50, 75, 90, 95, 99])
print(f"\nTRAP #2 - length (words): median={q[0]:.0f} p75={q[1]:.0f} p90={q[2]:.0f} "
f"p95={q[3]:.0f} p99={q[4]:.0f} max={lens.max()}")
print(f" ~{(lens > MAX_LEN*0.75).mean()*100:.1f}% of reviews exceed MAX_LEN={MAX_LEN} "
f"tokens (rough words->tokens factor 1.3). Section 9 measures what that costs.")
h_tr = {hashlib.md5(t.encode()).hexdigest() for t in raw["train"]["text"]}
h_te = {hashlib.md5(t.encode()).hexdigest() for t in raw["test"]["text"]}
print(f"\nTRAP #3 - leakage: {len(h_tr & h_te)} exact duplicate reviews across "
f"train/test; {len(raw['train'])-len(h_tr)} dupes inside train itself.")
def clean(t):
return t.replace("
", " ").replace("
", " ").strip()
plt.figure(figsize=(11, 3.2))
plt.subplot(1, 2, 1)
plt.hist(np.clip(lens, 0, 1000), bins=60)
plt.axvline(MAX_LEN, ls="--", color="k", label=f"MAX_LEN={MAX_LEN}")
plt.title("Review length (words, clipped at 1000)"); plt.legend()
plt.subplot(1, 2, 2)
plt.bar(["neg", "pos"], np.bincount(raw["train"]["label"]))
plt.title("Train class balance (perfectly balanced)")
plt.tight_layout(); plt.show()
We configure the Colab environment, install the required libraries, apply the PEFT–torchao compatibility fix, and set deterministic seeds for reproducible experiments. We load the Stanford IMDb dataset, shuffle and subsample the train and test splits, and inspect class balance, review-length distributions, duplicate leakage, and HTML artifacts. We also visualize review lengths and label frequencies so we understand the dataset structure before building any models.
Copy CodeCopiedUse a different Browser
print("\n" + "=" * 79 + "\n3. TF-IDF BASELINE\n" + "=" * 79)
Xtr = [clean(t) for t in train_ds["text"]]; ytr = np.array(train_ds["label"])
Xte = [clean(t) for t in eval_ds["text"]]; yte = np.array(eval_ds["label"])
t0 = time.time()
tfidf_clf = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_features=300_000,
sublinear_tf=True, strip_accents="unicode"),
LogisticRegression(C=8.0, max_iter=2000, n_jobs=-1),
)
tfidf_clf.fit(Xtr, ytr)
p_tfidf = tfidf_clf.predict_proba(Xte)[:, 1]
acc_tfidf = accuracy_score(yte, p_tfidf > 0.5)
auc_tfidf = roc_auc_score(yte, p_tfidf)
print(f"trained in {time.time()-t0:.1f}s -> acc={acc_tfidf:.4f} auc={auc_tfidf:.4f}")
vec, lr = tfidf_clf.steps[0][1], tfidf_clf.steps[1][1]
feats, coefs = np.array(vec.get_feature_names_out()), lr.coef_[0]
order = np.argsort(coefs)
print("\nmost NEGATIVE n-grams:", ", ".join(feats[order[:12]]))
print("most POSITIVE n-grams:", ", ".join(feats[order[-12:]][::-1]))
print("\n" + "=" * 79 + "\n4. LoRA FINE-TUNING\n" + "=" * 79)
tok = AutoTokenizer.from_pretrained(MODEL_NAME)
def tokenize(batch):
return tok([clean(t) for t in batch["text"]], truncation=True, max_length=MAX_LEN)
tr_tok = (train_ds.map(tokenize, batched=True, remove_columns=["text"])
.rename_column("label", "labels"))
ev_tok = (eval_ds.map(tokenize, batched=True, remove_columns=["text"])
.rename_column("label", "labels"))
base = AutoModelForSequenceClassification.from_pretrained(
MODEL_NAME, num_labels=2,
id2label={0: "NEGATIVE", 1: "POSITIVE"},
label2id={"NEGATIVE": 0, "POSITIVE": 1},
)
lora_cfg = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_lin", "v_lin"],
modules_to_save=["pre_classifier", "classifier"],
)
try:
model = get_peft_model(base, lora_cfg)
except ImportError as e:
_disable_torchao_probe()
print(f"[compat] retrying after backend probe failure: {e}")
model = get_peft_model(base, lora_cfg)
model.print_trainable_parameters()
def compute_metrics(eval_pred):
logits, labels = eval_pred
probs = torch.softmax(torch.tensor(logits), dim=-1).numpy()[:, 1]
preds = (probs > 0.5).astype(int)
return {"accuracy": accuracy_score(labels, preds),
"f1_macro": f1_score(labels, preds, average="macro"),
"roc_auc": roc_auc_score(labels, probs)}
_ta = inspect.signature(TrainingArguments.__init__).parameters
_eval_key = "eval_strategy" if "eval_strategy" in _ta else "evaluation_strategy"
ta_kwargs = dict(
output_dir="./imdb_lora", learning_rate=LR,
per_device_train_batch_size=BATCH, per_device_eval_batch_size=BATCH * 2,
num_train_epochs=EPOCHS, weight_decay=0.01, warmup_ratio=0.06,
logging_steps=50, save_strategy="epoch", save_total_limit=1,
load_best_model_at_end=True, metric_for_best_model="accuracy",
fp16=(DEVICE == "cuda"), report_to="none", seed=SEED,
)
ta_kwargs[_eval_key] = "epoch"
_tr = inspect.signature(Trainer.__init__).parameters
_tok_key = "processing_class" if "processing_class" in _tr else "tokenizer"
trainer = Trainer(
model=model, args=TrainingArguments(**ta_kwargs),
train_dataset=tr_tok, eval_dataset=ev_tok,
data_collator=DataCollatorWithPadding(tok),
compute_metrics=compute_metrics,
callbacks=[EarlyStoppingCallback(early_stopping_patience=2)],
**{_tok_key: tok},
)
t0 = time.time()
trainer.train()
print(f"\nfine-tuned in {(time.time()-t0)/60:.1f} min")
We train a strong TF-IDF and Logistic Regression baseline and inspect the most influential positive and negative n-grams to establish an interpretable reference point. We then tokenize the IMDb reviews and configure DistilBERT with LoRA adapters that update only a small subset of model parameters while keeping the backbone largely frozen. We use the Hugging Face Trainer with dynamic padding, early stopping, mixed precision, and multiple evaluation metrics to fine-tune the transformer efficiently.
Copy CodeCopiedUse a different Browser
print("\n" + "=" * 79 + "\n5. EVALUATION\n" + "=" * 79)
pred_out = trainer.predict(ev_tok)
p_lora = torch.softmax(torch.tensor(pred_out.predictions), dim=-1).numpy()[:, 1]
y_true = np.array(pred_out.label_ids)
yhat = (p_lora > 0.5).astype(int)
print(classification_report(y_true, yhat, target_names=["neg", "pos"], digits=4))
cm = confusion_matrix(y_true, yhat)
fig, ax = plt.subplots(1, 2, figsize=(11, 4))
ax[0].imshow(cm, cmap="Blues")
for i in range(2):
for j in range(2):
ax[0].text(j, i, cm[i, j], ha="center", va="center", fontsize=14)
ax[0].set_xticks([0, 1], ["pred neg", "pred pos"])
ax[0].set_yticks([0, 1], ["true neg", "true pos"]); ax[0].set_title("Confusion matrix")
for name, p in [("TF-IDF", p_tfidf), ("DistilBERT+LoRA", p_lora)]:
fpr, tpr, _ = roc_curve(y_true, p)
ax[1].plot(fpr, tpr, label=f"{name} (AUC={roc_auc_score(y_true, p):.4f})")
ax[1].plot([0, 1], [0, 1], "k--", lw=0.8)
ax[1].set_xlabel("FPR"); ax[1].set_ylabel("TPR"); ax[1].set_title("ROC"); ax[1].legend()
plt.tight_layout(); plt.show()
print("\n" + "=" * 79 + "\n6. THRESHOLD & CALIBRATION\n" + "=" * 79)
ths = np.linspace(0.05, 0.95, 91)
accs = [(y_true == (p_lora > t)).mean() for t in ths]
best_t = ths[int(np.argmax(accs))]
print(f"acc@0.50 = {accs[45]:.4f} | best threshold = {best_t:.2f} -> acc = {max(accs):.4f}")
def expected_calibration_error(probs, labels, n_bins=10):
"""ECE: |confidence - accuracy| averaged over confidence bins."""
conf = np.maximum(probs, 1 - probs)
correct = (probs > 0.5).astype(int) == labels
bins = np.linspace(0, 1, n_bins + 1)
ece, xs, ys = 0.0, [], []
for lo, hi in zip(bins[:-1], bins[1:]):
m = (conf > lo) & (conf 0.5).astype(int)
err["correct"] = err.pred == err.y
err["confidence"] = np.maximum(err.p_pos, 1 - err.p_pos)
print("--- 3 most CONFIDENT mistakes (where the model is confidently wrong) ---")
for _, r in err[~err.correct].nlargest(3, "confidence").iterrows():
print(f"\n[true={'pos' if r.y else 'neg'} pred={'pos' if r.pred else 'neg'} "
f"conf={r.confidence:.3f} words={r.n_words}]")
print(clean(r.text)[:400].replace("\n", " "), "...")
err["bucket"] = pd.qcut(err.n_words, 4, labels=["short", "med", "long", "v.long"])
by_len = err.groupby("bucket", observed=True).agg(acc=("correct", "mean"), n=("correct", "size"))
print("\n--- accuracy by review length (truncation hurts long reviews) ---")
print(by_len.to_string())
print("\n" + "=" * 79 + "\n8. OCCLUSION SALIENCY\n" + "=" * 79)
infer_model = model.merge_and_unload()
infer_model.to(DEVICE).eval()
@torch.no_grad()
def predict_proba(texts, bs=64):
out = []
for i in range(0, len(texts), bs):
enc = tok([clean(t) for t in texts[i:i + bs]], truncation=True,
max_length=MAX_LEN, padding=True, return_tensors="pt").to(DEVICE)
out.append(torch.softmax(infer_model(**enc).logits, dim=-1)[:, 1].cpu().numpy())
return np.concatenate(out)
def occlusion(text, max_words=60):
words = clean(text).split()[:max_words]
base = predict_proba([" ".join(words)])[0]
variants = [" ".join(words[:i] + words[i + 1:]) for i in range(len(words))]
dropped = predict_proba(variants)
return words, base - dropped, base
sample = err[err.correct].nlargest(1, "confidence").iloc[0]
words, contrib, base_p = occlusion(sample.text)
print(f"P(positive) for the full excerpt = {base_p:.3f} "
f"(true label = {'pos' if sample.y else 'neg'})\n")
top = np.argsort(np.abs(contrib))[-15:]
plt.figure(figsize=(7, 5))
plt.barh(range(len(top)), contrib[top],
color=["tab:green" if contrib[i] > 0 else "tab:red" for i in top])
plt.yticks(range(len(top)), [words[i] for i in top])
plt.xlabel("Δ P(positive) when the word is removed")
plt.title("Occlusion saliency — green pushes POSITIVE, red pushes NEGATIVE")
plt.tight_layout(); plt.show()
print("\n" + "=" * 79 + "\n9. HEAD vs TAIL TRUNCATION\n" + "=" * 79)
probe = err.nlargest(600, "n_words")
W = 180
head_txt = [" ".join(clean(t).split()[:W]) for t in probe.text]
tail_txt = [" ".join(clean(t).split()[-W:]) for t in probe.text]
yp = probe.y.values
acc_head = ((predict_proba(head_txt) > 0.5).astype(int) == yp).mean()
acc_tail = ((predict_proba(tail_txt) > 0.5).astype(int) == yp).mean()
print(f"on the {len(probe)} longest reviews, using only {W} words:")
print(f" first {W} words -> acc {acc_head:.4f}")
print(f" last {W} words -> acc {acc_tail:.4f}")
print(" Practical takeaway: if the tail wins, feed head+tail to the model or "
"raise MAX_LEN, rather than blindly truncating from the left.")
We examine the model’s most confident incorrect predictions and group reviews by length to identify truncation-related failure patterns and difficult examples. We merge the LoRA adapters into the underlying model and apply leave-one-word-out occlusion to estimate which words push individual predictions toward positive or negative sentiment. We then compare predictions based on the beginning and ending portions of long reviews to determine where the strongest sentiment information resides.
Copy CodeCopiedUse a different Browser
print("\n" + "=" * 79 + "\n10. PSEUDO-LABELLING\n" + "=" * 79)
unsup = raw["unsupervised"].shuffle(seed=SEED).select(range(N_UNSUP))
p_uns = predict_proba(unsup["text"])
keep = (p_uns > 0.95) | (p_uns 0.5).astype(int)
print(f"kept {keep.sum()}/{N_UNSUP} pseudo-labels at conf>0.95 "
f"(balance: {np.bincount(pl_labels)})")
aug = make_pipeline(
TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_features=300_000,
sublinear_tf=True, strip_accents="unicode"),
LogisticRegression(C=8.0, max_iter=2000, n_jobs=-1),
).fit(Xtr + pl_texts, np.concatenate([ytr, pl_labels]))
acc_aug = accuracy_score(yte, aug.predict(Xte))
print(f"TF-IDF baseline : {acc_tfidf:.4f}")
print(f"TF-IDF + pseudo-labels: {acc_aug:.4f} (Δ {acc_aug-acc_tfidf:+.4f})")
print("Caveat: gains are bounded by the teacher. Self-training also amplifies "
"the teacher's biases — always validate on clean, held-out data.")
print("\n" + "=" * 79 + "\n11. SAVE & INFER\n" + "=" * 79)
SAVE_DIR = "./imdb-distilbert-lora-merged"
infer_model.save_pretrained(SAVE_DIR); tok.save_pretrained(SAVE_DIR)
print(f"saved merged model to {SAVE_DIR}/ (load with "
f"AutoModelForSequenceClassification.from_pretrained('{SAVE_DIR}'))")
demos = [
"A masterclass in tension. The final act left the whole theatre silent.",
"Two hours I will never get back. Wooden acting, incoherent plot.",
"It's not the disaster the trailer promised, but it never really lands either.",
]
for d, p in zip(demos, predict_proba(demos)):
print(f" P(pos)={p:.3f} -> {'POSITIVE' if p > 0.5 else 'NEGATIVE'} | {d}")
print("\n" + "=" * 79)
print(f"SUMMARY (n_train={N_TRAIN}, n_eval={N_EVAL}, max_len={MAX_LEN})")
print("=" * 79)
print(pd.DataFrame([
{"model": "TF-IDF + LogReg", "accuracy": acc_tfidf, "roc_auc": auc_tfidf},
{"model": "TF-IDF + pseudo-labels", "accuracy": acc_aug, "roc_auc": float("nan")},
{"model": "DistilBERT + LoRA", "accuracy": accuracy_score(y_true, yhat),
"roc_auc": roc_auc_score(y_true, p_lora)},
]).to_string(index=False))
print("""
NEXT EXPERIMENTS
- Set FULL_RUN = True for the real 25k/25k benchmark (~40 min on a T4).
- Swap MODEL_NAME to 'roberta-base' (target_modules=['query','value']) or
'answerdotai/ModernBERT-base' for an 8k context window — no truncation.
- Head+tail truncation: first 128 + last 128 tokens, motivated by section 9.
- Ablate LoRA rank r in {4, 8, 16, 64} and plot accuracy vs trainable params.
- Replace the pseudo-label teacher with an ensemble and iterate self-training.
- Push to the Hub: huggingface_hub.login() then infer_model.push_to_hub(...).
""")
We use the fine-tuned transformer to generate high-confidence pseudo-labels for examples from IMDb’s unlabeled split and add these examples to the TF-IDF training corpus.
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み