视频分割
🕒 阅读时间:约 14 分钟
📅 2026-03 深度精选
Meta SAM 2 / SAM 2.1 视频与图像全景分割大模型:显存占用调优与生产流水线集成
Meta 的 Segment Anything Model 2 (SAM 2) 实现了从静态图片到动态视频画面的全景交互式分割。本文详解其时空记忆网络原理,并给出镜像拉取与视频流式跟踪流水线实操。
✂️ 一、SAM 2 时空记忆注意力架构
SAM 2 引入了流式记忆机制(Memory Attention):在视频处理过程中,模型将过去帧的交互点击与掩码(Mask)特征编码暂存至时空记忆队列,即使目标在视频中发生严重遮挡、形变或出画再入画,也能实现精准连续追踪。
⚡ 二、镜像拉取与视频分割跟踪调用
export HF_ENDPOINT="https://hf-mirror.net"
huggingface-cli download facebook/sam2.1-hiera-large \
--local-dir /data/models/sam2.1-hiera-large \
--local-dir-use-symlinks False
from sam2.build_sam import build_sam2_video_predictor
checkpoint = "/data/models/sam2.1-hiera-large/sam2.1_hiera_large.pt"
model_cfg = "configs/sam2.1/sam2.1_hiera_l.yaml"
predictor = build_sam2_video_predictor(model_cfg, checkpoint)
# 初始化视频预测状态
inference_state = predictor.init_state(video_path="input_video_frames/")
# 在第 0 帧点击提示目标
_, out_obj_ids, out_mask_logits = predictor.add_new_points_or_box(
inference_state=inference_state,
frame_idx=0,
obj_id=1,
points=[[210, 350]],
labels=[1]
)
print("SAM 2 首帧提示成功,准备整段视频掩码传播!")