跳到正文
原文
Hugging Face Blog·· 11 天前精选AI 评分74

Transformers 支持直接加载并运行 llama.cpp 的 GGUF 量化模型

Transformers now runs llama.cpp quants

AI 导读

Transformers 新增对 GGUF 格式模型的直接加载支持,可通过 from_pretrained 配合 gguf_file 参数运行 Unsloth 等发布的 GGUF 量化检查点。

推荐理由

transformers 现已原生支持 GGUF 模型加载,通过复用 ggml Metal kernel 在 Apple Silicon 上获得接近 llama.cpp 的本地推理速度。

来源:Hugging Face Blog · huggingface.co