Open Weight · Quantization Explorer

Make models smaller. Understand the trade-offs.

A practical guide to model quantization — from FP16 and BF16 to FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF-based local inference.

FP88-bit floating-point workflows
INT8Lower-memory integer inference
INT4High compression for deployment
Trade-offsMemory · quality · speed · support
01 · Foundation

What is model quantization?

Quantization reduces the precision used to represent model weights or activations. The goal is usually to reduce memory requirements and make models easier or cheaper to run while preserving as much model quality as possible.

Hugging Face describes quantization as lowering model memory requirements by storing weights at lower precision while trying to preserve accuracy.
FP32 / BF16 / FP16
→
Quantization method
→
FP8 / INT8 / INT4
→
Lower memory
+
Potential speed gains
+
Trade-offs
02 · Precision

Bits per parameter: the basic intuition

The chart below shows theoretical raw weight storage relative to FP32. It ignores runtime overhead, metadata, KV cache, activations and mixed-precision components.

32FP32
16FP16/BF16
8FP8
8INT8
4INT4
Lower bit width does not guarantee faster inference. Actual performance depends on kernels, hardware, memory bandwidth, runtime support and the quantization method.
03 · Memory estimator

Estimate raw weight storage

Use this simple calculator to estimate the theoretical storage of model weights at a chosen bit width.

16.00 GB
Approximate decimal GB for raw parameters only. Real deployment memory can be higher.
04 · Main approaches

Quantization is not one technique

On-the-fly

Quantize during model loading rather than distributing a separately pre-quantized checkpoint.

bitsandbytes

Post-training

Quantize an already trained model, often using calibration or optimization to reduce error.

GPTQAWQ

Runtime ecosystem

Convert and quantize for a deployment stack such as GGUF / llama.cpp.

GGUFllama.cpp
05 · bitsandbytes

4-bit and 8-bit loading in Transformers

Hugging Face documents bitsandbytes as providing memory-efficient 8-bit and 4-bit linear layers and quantization integrations for Transformers.

LLM.int8()

An 8-bit method designed to preserve higher precision for sensitive computations instead of naively forcing everything into INT8.

QLoRA

Uses 4-bit quantization with trainable low-rank adapter parameters, making parameter-efficient adaptation possible with a smaller memory footprint.

Hugging Face bitsandbytes documentation ↗

06 · GPTQ

Error-aware post-training quantization

Current Transformers documentation uses GPT-QModel as the maintained GPTQ backend. GPTQ is a post-training method that quantizes weight matrices while optimizing to reduce quantization error.

Hugging Face notes that current GPTQ workflows can quantize weights to low-bit representations such as INT4 and dequantize them during inference in optimized kernels.

Hugging Face GPTQ documentation ↗

07 · AWQ

Activation-aware weight quantization

AWQ focuses on preserving weights that are especially important to model behavior while compressing the model to low-bit representations.

Transformers documents AWQ as an activation-aware approach designed for 4-bit compression with limited performance degradation.

Hugging Face AWQ documentation ↗

08 · GGUF / llama.cpp

Quantization for local and portable inference

The llama.cpp ecosystem provides many quantized GGUF tensor types and tooling for converting higher-precision GGUF models into smaller quantized variants.

High-precision model ↓ Convert to GGUF ↓ llama-quantize ↓ Q8 / Q6 / Q5 / Q4 / lower-bit variants ↓ Evaluate quality + performance ↓ Run with llama.cpp

llama.cpp documents integer quantization from very low bit widths through 8-bit variants. The practical choice depends on model family, quality target and hardware.

llama.cpp quantization tools ↗

09 · Comparison

Common quantization directions

ApproachTypical bit widthStrengthWatch for
bitsandbytes4 / 8Convenient Transformers integration and on-the-fly loadingHardware/backend support and training limitations
GPTQCommonly 4; other bit widths supported by current backendsPost-training compression with error-aware optimizationKernel, model and checkpoint compatibility
AWQ4Activation-aware preservation of important weightsToolchain and runtime compatibility
GGUF / llama.cppMultiple low-bit typesStrong local-inference ecosystem and many quantization variantsModel architecture support and quality/runtime trade-offs
FP88Lower precision while remaining floating pointHardware and kernel support
10 · Decision helper

Which direction should you investigate?

Select a deployment goal. This is a starting point, not a universal recommendation.

Start by evaluating bitsandbytes.

Its 4-bit and 8-bit Transformers integration makes it a practical entry point when your model and hardware are supported.

11 · Trade-offs

What should you measure?

Memory

How much RAM or VRAM is actually required?

Quality

How much task performance changes after quantization?

Latency

Does the runtime and hardware actually become faster?

Compatibility

Can your serving stack load and accelerate the chosen format?

Always benchmark on the real workload. A smaller checkpoint can still perform worse operationally if the runtime lacks optimized kernels for that quantization.
12 · Common mistakes

Quantization misconceptions

Bits ≠ method

Two 4-bit methods can behave very differently.

Smaller ≠ faster

Speed depends on kernels, hardware and runtime support.

Format ≠ quantization

GGUF or Safetensors are serialization formats; quantization describes numerical representation and method.

Memory ≠ file size only

KV cache, activations and runtime overhead also matter.

Quality loss is task-specific

Benchmark the model on the tasks that matter to you.

Support changes

Quantization libraries and hardware backends evolve quickly.

13 · Quick checklist

Before choosing a quantization

QuestionWhy it matters
What hardware will run the model?Backend support and optimized kernels differ by platform.
Which runtime will serve it?Not every runtime supports every quantization method.
What memory limit do you have?Defines how aggressive compression may need to be.
What quality loss is acceptable?Lower bit widths can affect downstream performance.
Do you need fine-tuning?Some workflows support PEFT or adapter training better than others.
Do you need portability?A highly optimized method may tie you to a specific runtime or hardware stack.
Next

Continue the Open Weight series

Open Weight Explorer

Understand tensors, model weights and the deployment stack.

Weight Format Explorer

Safetensors, GGUF, metadata, sharding and conversion.

Model Portability Explorer

Formats, runtimes, hardware and compatibility.

Primary sources

Technical references

Hugging Face Transformers — Quantization overview ↗

Hugging Face — bitsandbytes ↗

Hugging Face — GPTQ ↗

Hugging Face — AWQ ↗

llama.cpp — quantization tools ↗