English
← Back to AI Technology

Inference and deployment

Continuous Batching / PagedAttention

Serving techniques that continuously admit and retire requests while managing KV-cache memory in pages. Together they improve utilization and throughput under variable-length workloads.

Reference period
2022–2023
Last reviewed

Key terms

  • request scheduling
  • memory paging
  • throughput
  • serving

Connected across the project

Primary sources

  1. Orca: A Distributed Serving System for Transformer-Based Generative Models
  2. Efficient Memory Management for Large Language Model Serving with PagedAttention

This is a curated technical map, not a claim of comprehensive coverage.