한국어
← AI 기술 체계로 돌아가기

추론과 배포

Continuous Batching / PagedAttention

요청을 계속 받아들이고 완료하면서 KV 캐시 메모리를 페이지 단위로 관리하는 서빙 기술이다. 가변 길이 워크로드에서 활용률과 처리량을 높인다.

기준 시기
2022–2023
최근 검토

핵심 용어

  • request scheduling
  • memory paging
  • throughput
  • serving

프로젝트 내 연결

1차 출처

  1. Orca: A Distributed Serving System for Transformer-Based Generative Models
  2. Efficient Memory Management for Large Language Model Serving with PagedAttention

선별한 기술 지도이며 AI 기술을 완전히 망라한다는 뜻이 아니다.