Conceptual

LLM Weight Quantization

Running large models in 8- and 4-bit: post-training quantization, outlier channels, and calibration-based methods like GPTQ and AWQ.