Inference and deployment
Open-model Inference Ecosystem
Engines and runtimes such as vLLM and llama.cpp turn downloadable model weights into local or server inference. Their priorities differ across throughput, hardware portability, model formats, and operational complexity.
Key terms
- vLLM
- llama.cpp
- GGUF
- local inference
Connected across the project
Related technology
Historical context
Related sections
Primary sources
This is a curated technical map, not a claim of comprehensive coverage.