hf_text-generation-inference

Commit Graph

Author	SHA1	Message	Date
Daniël de Kok	959b9dc25f	Fixup constructor arguments	2024-07-17 07:42:24 +00:00
Daniël de Kok	27ef5aa029	Sync allocator interfaces	2024-07-16 14:42:32 +00:00
Daniël de Kok	48b21eab7a	Last accessed fixes	2024-07-16 11:55:46 +00:00
Daniël de Kok	dd2d6cfe40	Proper support for two allocations with overlapping prefixes	2024-07-16 11:40:35 +00:00
Daniël de Kok	d4ce5389ce	Add another problematic case	2024-07-16 10:20:11 +00:00
Daniël de Kok	0e6ff1293a	Fixes	2024-07-16 10:10:10 +00:00
Daniël de Kok	7c046c9190	First step towards cleaning up Breaks tests, but I want to shuffle around data structures so that we can just pass block ids to free.	2024-07-15 14:14:10 +00:00
Daniël de Kok	05611f6b40	Renaming, window size	2024-07-15 13:58:10 +00:00
Daniël de Kok	083806aa42	Traitify the current allocator in preparation for swappable alloc	2024-07-15 13:44:22 +00:00
Daniël de Kok	3b4754cd31	Better leaf tracking	2024-07-12 16:03:21 +02:00
Daniël de Kok	1a461234d5	Avoid continuous sorting during reclamation	2024-07-12 13:39:22 +02:00
Daniël de Kok	c352a3e231	Shake out some issues, add correct removal order test	2024-07-12 13:39:22 +02:00
Daniël de Kok	6d0094e5d4	docs/cleanups	2024-07-12 13:39:22 +02:00
Daniël de Kok	3b6bef4078	Walk up to predecessors	2024-07-12 13:39:22 +02:00
Daniël de Kok	9da64a7b16	Basic test passes	2024-07-12 13:39:22 +02:00
Daniël de Kok	dbb82e274c	WIP	2024-07-12 13:39:22 +02:00
Daniël de Kok	dbb23fbfa8	Use symmetric quantization in the `quantize` subcommand (#2120 ) Packing of asymmetric quantization is broken, all (q)zeros values of `0` get reset to `1`, resulting in a loss of accuracy. So instead use symmetric quantization. To be able to distinguish models with symmetric and asymmetric quantization, a new config tensor `gptq_sym` is added. If this tensor is not present, we assume `sym=False`.	2024-07-12 12:20:12 +02:00
SeongBeomLEE	c46eaf707b	[fix] Modifying base in yarn embedding (#2212 )	2024-07-12 10:04:51 +02:00
drbh	d789de329a	fix: append DONE message to chat stream (#2221 ) * fix: append DONE message to chat stream * fix: update completions endpoint	2024-07-11 10:42:58 -04:00
Daniël de Kok	cb150eb295	Add support for FP8 on compute capability >=8.0, <8.9 (#2213 ) Use FP8 GPTQ-Marlin kernels to enable FP8 support on CUDA GPUs with compute capability >=8.0 and <8.9. Co-authored-by: Florian Zimmermeister <flozi00.fz@gmail.com>	2024-07-11 16:03:26 +02:00
Daniël de Kok	8511669cb2	Move quantized weight handling out of the `Weights` class (#2194 ) Quantized weights were loaded in the `Weights` class, but this was getting quite unwieldy, where every higher level method to load weights was a long conditional to cover all the different quantizers. This change moves loading of quantized weights out of the `Weights` class. This is done by defining a simple `WeightsLoader` interface that is implemented by `Exl2WeightsLoader`, `GPTQWeightsLoader`, and `MarlinWeightsLoader`. These implementations are in the quantizers' respective modules. The `Weights` class provides the low-level load operations (such as loading tensors or sharded tensors), but delegates loads that need quantizer-specific weight processing to a loader. The loaders still use the low-level functionality provided by `Weights`. I initially tried making a hierarchy where a class like `GPTQWeights` would inherit from `Weights`. But it is not very flexible (e.g. does not work well with the new weight storage mock used in tests) and the implicit indirections made the code harder to follow.	2024-07-09 20:04:03 +02:00
Nicolas Patry	4c976fb406	Updating the self check (#2209 ) * Updating the self check * Fix. * Revert the CLI . * cli. * Space. * Revert cargo update.	2024-07-09 17:23:48 +02:00
vinkamath	f5ba9bfd52	Fixed README ToC (#2196 ) Co-authored-by: Vinayak Kamath <Vinayak.Kamath@target.com>	2024-07-09 11:22:08 +02:00
Nicolas Patry	fe710af25f	Adding sanity check to openapi docs.	2024-07-09 11:13:48 +02:00
Guillaume LEGENDRE	5e2a305880	Fix buildx cache + change runner type (#2176 ) * Update build.yaml * Update build.yaml * change to S3 cache * change to CPU Runners * remove comments	2024-07-08 18:13:32 +02:00
fxmarty	4c50b6d04b	Fix nccl regression on PyTorch 2.3 upgrade (#2099 ) * fix nccl issue * add note in dockerfile * use v2.22.3 that also fixes @samsamoa's repro * poetry actually can't handle the conflict between torch and nccl * set LD_PRELOAD	2024-07-08 17:52:10 +02:00
drbh	87ebb6477b	feat: use model name as adapter id in chat endpoints (#2128 )	2024-07-08 16:06:49 +02:00
Wang, Yi	58effe78b5	update to metrics 0.23.0 or could work with metrics-exporter-promethe… (#2190 ) update to metrics 0.23.0 or could work with metrics-exporter-prometheus 0.15.1 Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>	2024-07-08 16:03:59 +02:00
Javier Martinez	16d9e505fd	fix: python deserialization (#2178 )	2024-07-08 15:59:16 +02:00
Wang, Yi	07e240ca37	add doc for intel gpus (#2181 ) Signed-off-by: Wang, Yi A <yi.a.wang@intel.com>	2024-07-08 15:57:06 +02:00
Daniël de Kok	5c7c9f1390	Falcon/DBRX: get correct number of key-value heads (#2205 )	2024-07-08 13:22:38 +02:00
Daniël de Kok	153fcf7739	Fix incorrect cache allocation with multi-query (#2203 ) We wouldn't allocate any memory in multi-query (1 KV head). Fixes Starcoder et al.	2024-07-08 11:19:48 +02:00
Daniël de Kok	cce475a949	hotfix: Fix number of KV heads (#2202 ) Fix number of KV heads	2024-07-08 09:52:12 +02:00
icyboy™	521d0d990f	fix dbrx & opt model prefix bug (#2201 ) * Update idefics_causal_lm.py Fix syntax issues * fix dbrx & opt model prefix bug	2024-07-08 09:01:14 +02:00
Daniël de Kok	05c094fcfa	Consistently take `prefix` in model constructors (#2191 ) * Consistently take `prefix` in model constructors * Release test check fix * Misc refactor-related fixes	2024-07-05 16:07:48 +02:00
Daniël de Kok	67ef0649cf	GPTQ CI improvements (#2151 ) * Add more representative Llama GPTQ test The Llama GPTQ test is updated to use a model with the commonly-used quantizer config format and activation sorting. The old test is kept around (but renamed) since it tests the format produced by `text-generation-server quantize`. * Add support for manually triggering a release build	2024-07-05 14:12:16 +02:00
Daniël de Kok	b67d46336e	Fix Starcoder2 after refactor (#2189 )	2024-07-05 12:22:45 +02:00
Nicolas Patry	853d4eb9cf	Hotfixing after refactor.	2024-07-05 09:25:29 +00:00
Nicolas Patry	fb2f74e2b9	Refactor dead code - Removing all `flash_xxx.py` files. (#2166 ) * Refactor dead code. * First working step. * Remove a lot of duplicated code. * More dead code. * More cleanup. * Fix Santacoder test. * Fixing the simple tests. * Fixing sharding. * Fixes for VLM. * Fixing santacoder (num_kv_heads hardcoded). * Removing more dead code. * Fixing `config.n_head`. * Stopping earlier because of `<end_of_utterance>` in idefics2. * Addresses comments. * Removing the dead code. * Fuse back mistral into FlashCausalLM. * Finish removal. * Fixing docs + causal_lm `batch_class`. * Fixing docs + causal.lm. * Add default to Gemma Causality. * Default value for gemma/gemma2. * Wrong default.	2024-07-05 10:29:56 +02:00
Aaron Mihalik	c6bcadf883	Adding "longrope" for Phi-3 (#2172 ) (#2179 ) Adding "longrope" for phi-3	2024-07-05 09:46:41 +02:00
Nicolas Patry	245d3de948	Preparing patch release. (#2186 )	2024-07-04 10:55:33 +02:00
Nicolas Patry	5ad41aa2a6	Fixing missing `object` field for regular completions. (#2175 ) * Fixing missing `object` field for regular completions. * Fixing docs by re-adding missing `Prompt`.	2024-07-03 12:56:27 +02:00
Nicolas Patry	2b3bd1e008	Fixing the dockerfile warnings. (#2173 )	2024-07-03 12:48:45 +02:00
Nicolas Patry	be4a4c47f9	Revert "Fixing missing `object` field for regular completions." This reverts commit `2bbb7fa4b2`.	2024-07-03 10:41:39 +00:00
Nicolas Patry	2bbb7fa4b2	Fixing missing `object` field for regular completions.	2024-07-03 10:40:22 +00:00
drbh	571530dd9a	feat: improve update_docs for openapi schema (#2169 ) * feat: add pre commit step to force schema update when router changes * fix: prefer improved update_doc and start server and compare * fix: adjust typo * fix: adjust revert typo * fix: update workflow to use update_doc md command * feat: improve workflow to check openapi schema too * fix: adjust timeout for CI * fix: adjust raise condition and install server in ci * fix: install protoc before server * feat: improve update doc and add command to print router schema * fix: adjust autodoc workflow * fix: explicitly install protoc and python * fix: alllow trailing space in openapi schema diff	2024-07-03 09:53:35 +02:00
Nicolas Patry	0759ec495e	Hotfixing qwen2 and starcoder2 (which also get clamping). (#2167 )	2024-07-02 14:26:47 +02:00
Guillaume LEGENDRE	963b6c6f0f	Ci test (#2124 ) * first test with registry mirror * change push registry * remove comments * Move cache to push registry * fix registry url * Update .github/workflows/ci_build.yaml --------- Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com>	2024-07-02 12:45:38 +02:00
Nicolas Patry	dea9c0dc74	Fixing rocm. (#2164 )	2024-07-02 12:01:08 +02:00
drbh	b966bc0d35	fix: use the base layers weight in mistral rocm (#2155 )	2024-07-02 11:56:25 +02:00

1 2 3 4 5 ...

870 Commits All Branches Search

870 Commits

All Branches