Hugging Face Blog·· 11 天前精选AI 评分74
Transformers 支持直接加载并运行 llama.cpp 的 GGUF 量化模型
Transformers now runs llama.cpp quants
AI 导读
Transformers 新增对 GGUF 格式模型的直接加载支持,可通过 from_pretrained 配合 gguf_file 参数运行 Unsloth 等发布的 GGUF 量化检查点。
推荐理由
transformers 现已原生支持 GGUF 模型加载,通过复用 ggml Metal kernel 在 Apple Silicon 上获得接近 llama.cpp 的本地推理速度。
来源:Hugging Face Blog · huggingface.co