History

Daniël de Kok 8deeaca4ff Add support for prefix caching to the v3 router (#2392 ) This change adds support for prefix caching to the v3 router. This is broken up from the backend support to ease reviewing. For now prefix caching is only enabled with `USE_PREFIX_CACHING=1` in this case, the router will switch to `RadixAllocator`. This allocator uses a radix trie to keep track of prefills that were seen prior. If a new prefill is a prefix of a previously-seen prefil, the router will send a request with `prefix_len>0`, which can be used by the backend to decide to reuse KV blocks from the cache, rather than recomputing them. Even though backend support is not added in this PR, the backend will still work with prefix caching enabled. The prefix lengths are just ignored and not used.		2024-08-12 14:59:17 +02:00
..
custom_kernels	chore: add pre-commit (#1569 )	2024-02-16 11:58:58 +01:00
exllama_kernels	MI300 compatibility (#1764 )	2024-05-17 15:30:47 +02:00
exllamav2_kernels	chore: add pre-commit (#1569 )	2024-02-16 11:58:58 +01:00
tests	feat: add ruff and resolve issue (#2262 )	2024-07-26 10:29:09 -04:00
text_generation_server	Add support for prefix caching to the v3 router (#2392 )	2024-08-12 14:59:17 +02:00
.gitignore	Impl simple mamba model (#1480 )	2024-02-08 10:19:45 +01:00
Makefile	hotfix: update nccl	2024-07-23 23:31:28 +02:00
Makefile-awq	chore: add pre-commit (#1569 )	2024-02-16 11:58:58 +01:00
Makefile-eetq	Upgrade EETQ (Fixes the cuda graphs). (#1729 )	2024-04-12 08:15:28 +02:00
Makefile-fbgemm	Upgrade fbgemm (#2398 )	2024-08-12 14:08:38 +02:00
Makefile-flash-att	Hotfixing `make install`. (#2008 )	2024-06-04 23:34:03 +02:00
Makefile-flash-att-v2	Softcapping for gemma2. (#2273 )	2024-07-22 18:27:10 +02:00
Makefile-lorax-punica	Enable multiple LoRa adapters (#2010 )	2024-06-25 14:46:27 -04:00
Makefile-selective-scan	chore: add pre-commit (#1569 )	2024-02-16 11:58:58 +01:00
Makefile-vllm	Add support for Deepseek V2 (#2224 )	2024-07-19 17:23:20 +02:00
README.md	chore: add pre-commit (#1569 )	2024-02-16 11:58:58 +01:00
poetry.lock	Install Marlin from standalone package (#2320 )	2024-07-29 15:37:10 +02:00
pyproject.toml	Install Marlin from standalone package (#2320 )	2024-07-29 15:37:10 +02:00
requirements_cuda.txt	hotfix: pin numpy (#2289 )	2024-07-23 17:53:19 +02:00
requirements_intel.txt	hotfix: pin numpy (#2289 )	2024-07-23 17:53:19 +02:00
requirements_rocm.txt	hotfix: pin numpy (#2289 )	2024-07-23 17:53:19 +02:00

README.md

Text Generation Inference Python gRPC Server

A Python gRPC server for Text Generation Inference

Install

make install

Run

make run-dev