English
← Back to AI Technology

Inference and deployment

Open-model Inference Ecosystem

Engines and runtimes such as vLLM and llama.cpp turn downloadable model weights into local or server inference. Their priorities differ across throughput, hardware portability, model formats, and operational complexity.

Reference period
2023–present
Last reviewed

Key terms

  • vLLM
  • llama.cpp
  • GGUF
  • local inference

Connected across the project

Primary sources

  1. vLLM documentation
  2. llama.cpp

This is a curated technical map, not a claim of comprehensive coverage.