← All writing
Paper Breakdown

GPTQ: what the post-training quantization paper actually says

Reading Frantar et al. (ETH Zürich, 2022) while trying to get a 30B fine-tuned model onto two GPUs.

The cost math was simple. We had a fine-tuned Falcon-30B checkpoint. At FP16, that's roughly 60 GB of weights. We were running two serving replicas for redundancy, needed a third for rolling deploys, and didn't have the budget to provision 12× A100 80GB cards. The obvious answer was quantization — cut weights to INT4, 4× compression, weights drop to ~15 GB, one A100 per replica. Done.

The naive quantizer destroyed quality. We ran our eval set — multi-turn instruction following with structured outputs — and the model that looked fine on perplexity benchmarks was silently mangling JSON schema boundaries and hallucinating tool invocations. We went from 94% schema compliance in FP16 to 71% in naive INT4.

This is when I actually read "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" (Frantar et al., ETH Zürich, 2022) instead of just running the library. The problem with naive quantization is specific, the fix is specific, and understanding both changes how you deploy and debug quantized models.

What naive quantization gets wrong

The simplest way to quantize a weight from FP16 to INT4: find the min and max of the weight tensor, divide that range into 16 intervals, and round each weight to the nearest integer. This is uniform quantization. It ignores one thing: not all weights matter equally to the model's output.

A linear layer computes y = Wx. The model's output is a function of Wx, not W alone. When you round weight w_ij to the nearest integer, the error you introduce propagates through the rest of the network scaled by the corresponding input activation. If certain inputs are consistently large across the calibration distribution, rounding errors in the weights they multiply get amplified. Weights connected to high-activation inputs are more important to preserve precisely.

The AWQ paper (Lin et al., MIT 2023) addresses this with per-channel scaling — identify the high-activation dimensions, upscale their weights before quantization so those weights snap to a finer grid. GPTQ addresses it differently: instead of preventing error on important weights, it compensates for the error after quantization by adjusting the other weights.

This compensation idea traces back to Optimal Brain Damage (LeCun et al., 1990) and Optimal Brain Surgeon (Hassibi and Stork, 1993). If you know the second-order structure of the loss — specifically, how the loss changes when you perturb each weight — you can quantize one weight and then update all remaining weights to partially cancel the error you just introduced. The second-order information comes from the Hessian.

For a linear layer, the Hessian simplifies. If you think of the quantization error as an additive perturbation to the weights, the second-order Taylor expansion of the output error is:

δE ≈ ½ × δW^T H δW

where H = 2 X^T X and X is the matrix of input activations from calibration data. This Hessian is computable without backpropagation — just forward passes through the layer.

Optimal Brain Quantization, and why it was too slow

The authors build on Optimal Brain Quantization (Hassibi et al., 2022), which applies OBS directly to the quantization problem. OBQ quantizes weights one at a time in an order determined by which quantization introduces the least immediate error. After quantizing each weight, it updates all remaining unquantized weights in the same row to compensate, using the OBS weight update formula:

δw_F = -(w_q - quant(w_q)) / [H_F^{-1}]_{qq} × [H_F^{-1}]_{:,q}

where q is the just-quantized weight index and F is the set of remaining weights. This is exact: you're projecting the quantization error onto the space spanned by the remaining weights.

The problem is computational cost. After each quantization, OBQ must re-invert a modified Hessian sub-matrix — a Schur complement update. For a weight matrix with d_col columns, this is O(d_col^3) work per weight, making the total cost O(d_col^4) per row. For GPT-class models with d_col = 4096, this is 4096^4 ≈ 2.8 × 10^14 operations per layer. OBQ quantizes a 1.75B-parameter model in roughly 2 hours. Scaling to 175B becomes ~200 hours per layer — obviously unusable.

The three optimizations that make GPTQ work at scale

GPTQ introduces three modifications that collectively reduce the computational cost to something tractable:

1. Arbitrary order quantization. OBQ's greedy ordering — quantize the weight that minimizes immediate error increase — requires computing a ranking criterion after every step. GPTQ observes empirically that for over-parameterized models (which most large language models are), the greedy ordering doesn't provide meaningful quality improvement over a fixed left-to-right column order. All rows of a weight matrix can be quantized in the same column order simultaneously. This allows the Hessian computation to be shared across rows, eliminating the per-row Hessian inversions.

2. Lazy batch updates. Rather than updating the full remaining weight matrix after each column is quantized, GPTQ processes columns in blocks of size B (B = 128 in the implementation). Within a block, the updates to out-of-block weights are accumulated but not applied. They're applied in one batched operation at the end of each block. This turns many small matrix-vector updates — bandwidth-bound on GPU — into one large matrix-matrix update, which is compute-bound and achieves much higher GPU utilization. The blocks can also be processed in parallel across rows.

3. Cholesky reformulation for numerical stability. The repeated Schur complement updates to H^{-1} as weights are quantized accumulate floating-point rounding errors. After hundreds of updates, the matrix can become non-positive-definite and produce nonsensical compensation deltas. GPTQ instead pre-computes the Cholesky decomposition of H (once, at the start) and uses it to read off specific rows of H^{-1} on demand. The Cholesky factors are numerically stable; you never explicitly form the full H^{-1} or apply repeated in-place updates. This replaces a numerically fragile iterative update with a single stable decomposition.

Together, these three changes bring the quantization cost for OPT-175B from the original OBQ estimate of hundreds of hours down to approximately four GPU-hours on a single A100 — the headline number in the paper title.

What the calibration data actually does

GPTQ doesn't require labeled data or gradient computation. Calibration uses 128 random samples from C4, the Colossal Clean Crawled Corpus — general web text. Each sample is truncated or padded to the model's context window. These samples are passed through each layer in sequence, and the resulting input activations X for each weight matrix are collected.

The Hessian is then H = 2 X^T X. That's it. No loss function, no gradients, no optimizer. The calibration data determines which input dimensions are "active" — large values in X correspond to weight columns where rounding errors will have large downstream effects.

Two implications:

  1. The calibration distribution matters. Using C4 (general web text) to quantize a model you're deploying for code generation means the Hessian is calibrated on text that doesn't resemble your inputs. The compensation computed may not adequately protect the weight dimensions that matter for your task.
  2. 128 samples is often enough. The paper shows diminishing returns beyond about 128 samples. But "enough" depends on the task distribution — if your task space is narrow and specific, calibrating on 128 domain-specific samples consistently outperforms 128 C4 samples.

What the performance numbers actually show

Quality results. For OPT-175B on WikiText-2 perplexity: FP16 baseline is approximately 8.34. GPTQ 4-bit is approximately 8.69 — a difference of 0.35 perplexity points, which is generally not distinguishable in downstream task performance. GPTQ 3-bit increases perplexity to roughly 10.9. Using grouped quantization (separate scale factors per group of 128 weights) at 3-bit pulls this down to approximately 9.3 — still noticeably worse than FP16 but usable for many tasks.

The degradation at 2-bit is more severe — perplexity roughly doubles — and 2-bit GPTQ is generally not deployment-ready without additional techniques.

Model size enabling. The paper's headline experiment: OPT-175B at FP16 requires ~350 GB of GPU memory. At INT4 GPTQ, this becomes ~87 GB — fitting on two A100 80GB cards with room for KV cache and activations. Before GPTQ, 175B-parameter inference required at minimum 8× A100s in a tensor-parallel configuration. GPTQ made it a two-card problem. The threshold for "what can you serve on commodity hardware" moved by 4×.

Inference speedup. The speedup from INT4 over FP16 comes from memory bandwidth savings, not compute savings — the dequantized weights still participate in FP16 matrix multiplications. The GPU's memory bandwidth is the bottleneck for autoregressive generation (decode phase), so cutting weight size cuts memory transfers. The paper reports approximately 3.25× throughput improvement on A100 and 4.5× on A6000 for OPT-175B at INT4. The A6000 numbers are higher because it has a less favorable memory bandwidth ratio than A100 — it's more bandwidth-starved in FP16, so INT4 helps more.

Production tradeoffs no one mentions in the benchmark post

Quantization takes hours and is sensitive to calibration. Running GPTQ on a 70B model on a single A100 takes 6-12 hours depending on context length and block size. This matters for pipelines that frequently update models — fine-tuned checkpoints, RLHF iterations, daily model refreshes. If you're doing active learning and updating weights weekly, the quantization time is in the critical path. AWQ (minutes to quantize) is often a better choice for update-heavy workflows even if GPTQ produces slightly better quality at the same bit-width.

The inference speedup requires a custom CUDA kernel. The 3.25× speedup in the paper assumes a custom INT4 matmul kernel (the paper describes their own; in practice, ExLlamaV2 and TensorRT-LLM provide production-quality versions). Stock PyTorch doesn't have optimized INT4 GEMM. If you run GPTQ weights with a dequantize-then-matmul approach — convert INT4 to FP16 on the fly for each operation — you get the memory savings but only a fraction of the throughput improvement. Many "I quantized to INT4 and it's only slightly faster" reports come from this.

INT4 inference is NVIDIA-only for production throughput. ExLlamaV2's optimized kernels target Ampere and Hopper architectures specifically. ROCm (AMD), Apple Silicon, and other accelerators have less mature INT4 GEMM support. If your serving infrastructure is heterogeneous or you're planning to move off NVIDIA, verify kernel availability before committing to GPTQ-quantized deployment.

Grouped quantization has memory overhead. The quality improvement from grouped quantization (separate scale/zero-point per 128 weights) comes at a cost: more scale parameters stored alongside the weights. For group size 128 on a 4096×4096 weight matrix, you have 4096×32 = 131,072 additional FP16 scale values per matrix. Small relative to the INT4 weights themselves but nonzero, and the overhead compounds across dozens of layers.

The quality gap widens on instruction-following and tool use. WikiText-2 perplexity is a weak proxy for production quality. Character-level instruction following, JSON schema compliance, function call formatting, and structured output tasks are more sensitive to quantization than next-word prediction on general text. The 0.35 perplexity gap at 4-bit often understates the actual quality difference on production tasks. Always evaluate on your actual eval set, not WikiText-2.

Failure modes in practice

The most common production failure: calibration distribution mismatch causing targeted capability degradation.

A team quantizes their fine-tuned coding model using 128 samples from C4 (the GPTQ library default). Perplexity on WikiText-2 looks fine. Code generation on HumanEval looks slightly degraded but within acceptable bounds. They ship it. Three weeks later, the on-call rotation starts seeing incidents where the model consistently fails on a specific pattern: multi-file context with import statements from multiple languages in the same prompt. The Hessian calibrated on English prose didn't adequately weight the dimensions important for code reasoning across language boundaries.

The fix: re-calibrate on a dataset that reflects your actual prompt distribution, even if it's just 128 samples. For code models, the CodeSearchNet corpus or your own production prompt sample. Recalibration is fast (the Hessian computation takes minutes); the Cholesky decomposition and quantization loop is what takes hours.

The second failure mode: silent quality regression at extreme sequence lengths. GPTQ's Hessian is computed at training-time context lengths. If you're serving at context lengths significantly longer than calibration — you calibrated at 2K, you're serving at 16K with RoPE extrapolation — the activation patterns that determine weight importance shift. Weights that were low-Hessian at 2K may be high-Hessian at 16K. This produces subtle degradation on long-context tasks that doesn't show up on short-context benchmarks.

When not to use GPTQ

When you need 8-bit quantization. LLM.int8() (bitsandbytes, Dettmers et al.) achieves 8-bit quantization with mixed-precision outlier handling in under a minute and requires no calibration data. For teams that want a quick reduction from FP16 memory without the calibration pipeline overhead, LLM.int8() is simpler. GPTQ at 4-bit offers more compression; LLM.int8() offers simpler deployment.

When you're fine-tuning the quantized model. GPTQ produces frozen quantized weights. If you want to fine-tune after quantization, use QLoRA instead — it quantizes to NF4 (a different quantization scheme optimized for normally-distributed weights) and adds trainable LoRA adapters on top. You can't run gradients through GPTQ-quantized weights in any useful way.

When your model update cycle is fast. If you're pushing new fine-tuned checkpoints every 24-48 hours, the 6-12 hour quantization loop becomes a bottleneck. Either parallelize quantization across multiple GPUs (each processing different layers), use AWQ instead (minutes), or run quantized and unquantized versions in a shadow configuration and promote only after quantization completes.

When you're on non-NVIDIA hardware. Without the ExLlamaV2/TensorRT-LLM kernel, INT4 throughput improvements are marginal. On AMD or Apple, GPTQ saves memory but may not provide meaningful latency improvement.

When task quality at 4-bit is unacceptable and 8-bit is sufficient. Measure before committing. 4-bit GPTQ saves 2× over 8-bit but introduces more quality degradation. If your eval shows 8-bit passes quality thresholds and 4-bit doesn't, there's no reason to use 4-bit. The memory difference between 8-bit and 4-bit may or may not be worth the quality cost depending on your serving constraints.

What the paper actually gives you

GPTQ demonstrated that second-order information is tractable for billion-parameter models if you engineer the computation carefully. The three optimizations — arbitrary order, lazy batching, Cholesky precomputation — are individually modest; together they produce a 10,000× speedup over the naive OBQ formulation. That engineering work is what made post-training quantization practical at GPT scale, not a new mathematical insight.

The broader framing: GPTQ treats quantization as a weight perturbation problem and uses second-order information to solve a least-squares compensation problem. AWQ treats quantization as a scaling problem and uses first-order activation information to reduce error before it happens. These are genuinely different approaches, and they tend to excel in different regimes — GPTQ at 3-bit and below (where the error to compensate is large), AWQ at 4-bit (where simple scaling is sufficient). In practice, at INT4, the two methods produce comparable quality and the choice usually comes down to tooling, quantization time, and target hardware.

For the serving problem that started this post: the Falcon-30B model quantized with domain-specific calibration data hit 91% schema compliance in INT4 — down from 94% in FP16 but above the 85% threshold we'd defined. We shipped on two A100s per replica. The C4-calibrated checkpoint was at 84%, below threshold, and would have caused a production incident. Calibration data was the variable that mattered.


GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers — Frantar, Ashkboos, Hoefler, Alistarh. ICLR 2023. arXiv:2210.17323.

Related reading

  • Mistral 7B: What the Sliding Window Attention Paper Actually Says

    A 7B model that outperforms Llama 2 13B everywhere and matches 34B on math. Mistral's sliding window attention and rolling KV cache aren't just architectural novelties — they're the answer to a specific memory problem that kills you at scale.

  • MLA: What the Multi-Head Latent Attention Paper Actually Says

    GQA cuts your KV cache 8x. At 128K context, you still run out of memory. Multi-Head Latent Attention — the architecture inside DeepSeek-V2, V3, and R1 — compresses KV cache via low-rank projection instead of head reduction, achieving 57x compression with near-zero quality loss. Here's what that actually means for inference.

  • BitNet b1.58: What the 1-bit LLM Paper Actually Says

    A 70B BitNet model fits in 7GB instead of 140GB — and the math says output quality matches FP16 at scale. The catch: you can't convert existing models. Here's what the paper actually proves, and why the hardware story matters more than the math.

← All writing