A practical guide to model quantization — from FP16 and BF16 to FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF-based local inference.
Quantization reduces the precision used to represent model weights or activations. The goal is usually to reduce memory requirements and make models easier or cheaper to run while preserving as much model quality as possible.
The chart below shows theoretical raw weight storage relative to FP32. It ignores runtime overhead, metadata, KV cache, activations and mixed-precision components.
Use this simple calculator to estimate the theoretical storage of model weights at a chosen bit width.
Quantize during model loading rather than distributing a separately pre-quantized checkpoint.
bitsandbytesQuantize an already trained model, often using calibration or optimization to reduce error.
GPTQAWQConvert and quantize for a deployment stack such as GGUF / llama.cpp.
GGUFllama.cppHugging Face documents bitsandbytes as providing memory-efficient 8-bit and 4-bit linear layers and quantization integrations for Transformers.
An 8-bit method designed to preserve higher precision for sensitive computations instead of naively forcing everything into INT8.
Uses 4-bit quantization with trainable low-rank adapter parameters, making parameter-efficient adaptation possible with a smaller memory footprint.
Current Transformers documentation uses GPT-QModel as the maintained GPTQ backend. GPTQ is a post-training method that quantizes weight matrices while optimizing to reduce quantization error.
AWQ focuses on preserving weights that are especially important to model behavior while compressing the model to low-bit representations.
Transformers documents AWQ as an activation-aware approach designed for 4-bit compression with limited performance degradation.
The llama.cpp ecosystem provides many quantized GGUF tensor types and tooling for converting higher-precision GGUF models into smaller quantized variants.
llama.cpp documents integer quantization from very low bit widths through 8-bit variants. The practical choice depends on model family, quality target and hardware.
| Approach | Typical bit width | Strength | Watch for |
|---|---|---|---|
| bitsandbytes | 4 / 8 | Convenient Transformers integration and on-the-fly loading | Hardware/backend support and training limitations |
| GPTQ | Commonly 4; other bit widths supported by current backends | Post-training compression with error-aware optimization | Kernel, model and checkpoint compatibility |
| AWQ | 4 | Activation-aware preservation of important weights | Toolchain and runtime compatibility |
| GGUF / llama.cpp | Multiple low-bit types | Strong local-inference ecosystem and many quantization variants | Model architecture support and quality/runtime trade-offs |
| FP8 | 8 | Lower precision while remaining floating point | Hardware and kernel support |
Select a deployment goal. This is a starting point, not a universal recommendation.
Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.
How much RAM or VRAM is actually required?
How much task performance changes after quantization?
Does the runtime and hardware actually become faster?
Can your serving stack load and accelerate the chosen format?
Two 4-bit methods can behave very differently.
Speed depends on kernels, hardware and runtime support.
GGUF or Safetensors are serialization formats; quantization describes numerical representation and method.
KV cache, activations and runtime overhead also matter.
Benchmark the model on the tasks that matter to you.
Quantization libraries and hardware backends evolve quickly.
| Question | Why it matters |
|---|---|
| What hardware will run the model? | Backend support and optimized kernels differ by platform. |
| Which runtime will serve it? | Not every runtime supports every quantization method. |
| What memory limit do you have? | Defines how aggressive compression may need to be. |
| What quality loss is acceptable? | Lower bit widths can affect downstream performance. |
| Do you need fine-tuning? | Some workflows support PEFT or adapter training better than others. |
| Do you need portability? | A highly optimized method may tie you to a specific runtime or hardware stack. |
Understand tensors, model weights and the deployment stack.
Safetensors, GGUF, metadata, sharding and conversion.
Formats, runtimes, hardware and compatibility.