AI Briefing
Sign in

Hugging Face Adds Native GGUF Execution to Transformers

·2026.09.29 11:30

Key point

Hugging Face Transformers has released a feature that allows direct inference of GGUF models in a compressed state without dequantization on Apple Silicon Macs, utilizing llama.cpp's ggml Metal kernels.

1 / 2

Details

Hugging Face has released a feature in the Transformers library that loads and runs GGUF weights in a compressed state without dequantization. Previously, loading GGUF files involved converting them to PyTorch dense models, which eliminated memory savings. This update reuses llama.cpp's ggml Metal kernels to efficiently manage GPU memory.

Key Features and Supported Scope

  • First Supported Targets: Apple Silicon Macs (MPS backend) and the Qwen3.8 architecture compatible with the Qwen3.5 series (Dense/MoE)
  • How It Works: Weights are loaded onto the GPU in their packed state, invoked via the gguf_file argument in AutoModelForCausalLM.from_pretrained
  • Kernel Utilization: Uses 5 Metal kernels including ggml-quantization, ggml-norm, ggml-attn, ggml-gated-delta-net, and topk to accelerate operations
  • Compatibility: Falls back to the existing dequantization method if compatible kernels are unavailable; other architectures like Llama and Mistral still use the dequantization loader

Performance Benchmarks

In a MacBook Pro M2 Max (32GB) environment compared to llama.cpp, the results demonstrated performance close to the C++ runtime despite being Python/PyTorch-based.

  • Qwen3.5-4B (Q4_K_M): Transformers 70.4 tok/s vs llama.cpp 71.8 tok/s (approx. 98% level)
  • Qwen3.8-27B: Transformers 15.9 tok/s vs llama.cpp 13.4 tok/s (Transformers superior)
  • Qwen3.5-35B-A3B (MoE): Transformers 60.2 tok/s vs llama.cpp 61.3 tok/s (approx. 98% level)

The performance gain from applying layer kernels is significant. The Qwen3.5-35B-A3B model increased from 28.8 tok/s using only the ggml-quantization kernel to 60.2 tok/s (2.09x) with all kernels applied.

Use Cases and Limitations

  • Use Cases: Experimenting with GGUF models in Python/PyTorch environments, prototyping custom layers, evaluating quality between original and quantized models, and fine-tuning (using the dequantize=True option)
  • Limitations: Currently Apple Silicon (MPS) only, with CUDA/CPU using the dequantization path. Batch processing with padding requires performance improvements, and the Hugging Face team plans to extend this to MPS's generate_batch. Currently, only the Qwen3.5/3.8 series is supported.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.