Inference and deployment
Continuous Batching / PagedAttention
Serving techniques that continuously admit and retire requests while managing KV-cache memory in pages. Together they improve utilization and throughput under variable-length workloads.
Key terms
- request scheduling
- memory paging
- throughput
- serving
Connected across the project
Builds on
Related technology
Historical context
Primary sources
This is a curated technical map, not a claim of comprehensive coverage.