Inference and deployment
Quantization
Representing weights, activations, or caches with lower numerical precision. It can reduce memory, bandwidth, and compute costs, with accuracy and hardware-support tradeoffs.
Key terms
- precision
- INT8
- INT4
- weight-only
Connected across the project
Related technology
Related papers
Historical context
Primary sources
This is a curated technical map, not a claim of comprehensive coverage.