🤗
HF-Mirror Engineering Guide
长文本基座 🕒 阅读时间:约 14 分钟 📅 2026-03 深度精选

智谱开源 GLM-4-9B-Chat 全能基座:1M 上下文长文本推理与函数调用 (Tool Call) 实战

智谱 AI 开源的 GLM-4-9B-Chat 支持百万级(1M Tokens)超长上下文,大海捞针检索准确率高达 100%。本文详解其长文本位置编码优化与企业级 Function Calling 工具调用落地实操。

📜 一、GLM-4-9B-Chat 的 1M 上下文与综合能力

GLM-4-9B-Chat 凭借 90 亿紧凑参数,全面超越了许多旧款 13B~34B 模型。原生支持长达 1,000,000 Tokens 的长文档输入(相当于两部《红楼梦》全文),并在复杂结构化数据解析与工具外部调用(Tool Use)上展现出优异的指令遵循率。

⚡ 二、镜像拉取与 Function Calling 结构化输出

下载 GLM-4-9B-Chat 权重
export HF_ENDPOINT="https://hf-mirror.net"
huggingface-cli download THUDM/glm-4-9b-chat \
  --local-dir /data/models/glm-4-9b-chat \
  --local-dir-use-symlinks False
tool_call_demo.py: 工具函数调用实操
from transformers import AutoModelForCausalLM, AutoTokenizer

path = "/data/models/glm-4-9b-chat"
tokenizer = AutoTokenizer.from_pretrained(path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(path, torch_dtype="auto", device_map="auto", trust_remote_code=True)

# 传入包含自定义工具定义的 Prompt
tools = [{
    "name": "get_stock_price",
    "description": "获取股票当前市场行情报价",
    "parameters": {"type": "object", "properties": {"symbol": {"type": "string"}}, "required": ["symbol"]}
}]

inputs = tokenizer.apply_chat_template([
    {"role": "user", "content": "请帮我查一下腾讯控股 (00700.HK) 现在的实时股价。"}
], tools=tools, return_tensors="pt", add_generation_prompt=True).to("cuda")

output = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(output[0]))
← 专栏目录 查看全部 32 篇开源大模型工程实录 动手配置 → 5分钟零门槛配置与高速拉取教程