🤗
HF-Mirror Engineering Guide
架构剖析 🕒 阅读时间:约 13 分钟 📅 2026-03 深度精选

Google Gemma 2 架构剖析:滑动窗口注意力 (SWA)、Logit 软截断与本地微调实战

Google DeepMind Gemma 2 引入了滑动窗口注意力与 Logit 软截断等多项学术前沿成果。本文剖析其数学原理,并提供基于镜像源的单卡低显存 LoRA 微调实战。

💎 一、Google Gemma 2 核心设计哲学

  • 交替滑动窗口注意力 (SWA): 偶数层使用 4096 窗口局部注意力,奇数层使用全局注意力,将长文本 KV 显存减少近半。
  • Logit Soft-Capping (软截断): 在注意力投影后使用 tanh 将激活值缩放截断至指定范围,防止极端梯度导致生成崩溃。

⚡ 二、镜像拉取与 Unsloth 极速微调

下载 Gemma 2 9B 权重
export HF_ENDPOINT="https://hf-mirror.net"
huggingface-cli download google/gemma-2-9b-it \
  --local-dir /data/models/gemma-2-9b-it \
  --local-dir-use-symlinks False
微调示例代码
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="/data/models/gemma-2-9b-it",
    max_seq_length=4096,
    load_in_4bit=True
)

model = FastLanguageModel.get_peft_model(
    model,
    r=16,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
    lora_alpha=16
)
print("Gemma 2 LoRA 适配器就绪!")
← 专栏目录 查看全部 32 篇开源大模型工程实录 动手配置 → 5分钟零门槛配置与高速拉取教程