preemo_text-generation-inference

Commit Graph

Author	SHA1	Message	Date
Michael Feil	012c917b6f	Wrapping completions and chat/completions endpoint (#2 ) * rebase and squash commits on latest main * cargo fmt * fix: 2038y problem --------- Co-authored-by: michaelfeil <me@michaelfeil.eu>	2023-09-27 08:58:07 -07:00
OlivierDehaene	1da642bd0e	feat(server): add local prom and health routes if running w/ ngrok	2023-07-21 16:56:30 +02:00
OlivierDehaene	b66b190403	feat(router): ngrok edge (#642 )	2023-07-19 11:59:58 +02:00
OlivierDehaene	b4024edd45	feat: better errors for warmup and TP (#575 ) Close #571	2023-07-10 14:47:15 +02:00
OlivierDehaene	e28a809004	v0.9.0 (#525 )	2023-07-01 19:25:41 +02:00
OlivierDehaene	e74bd41e0f	feat(server): add paged attention to flash models (#516 ) Closes #478	2023-06-30 19:09:59 +02:00
Robert Kimball	70f485bf9f	feat(router): add header option to disable buffering for the generate_stream response (#498 ) # This PR adds an http header option to disable buffering for the generate_stream endpoint response stream. Problem: If a model is run behind a proxy server such as nginx that has buffering enabled then the response stream from generate_stream gets aggregated into a single response which basically disables streaming. Instead of getting a chunked response where each token is presented over time the response presents everything all at once. Solution: This change adds the `X-Accel-Buffering` http header which disables buffering for the generate_stream response, allowing the response to stream properly.	2023-06-28 11:50:12 +02:00
OlivierDehaene	f59fb8b630	feat(router): add ngrok integration (#453 )	2023-06-16 16:25:11 +02:00
OlivierDehaene	895c5f1562	feat(server): only compute prefill logprobs when asked (#406 ) Close #288	2023-06-02 17:12:30 +02:00
OlivierDehaene	942005386a	feat(router): log input/ouput at debug level (#364 ) @njhill FYI	2023-05-23 20:47:37 +02:00
OlivierDehaene	e250282213	feat(docker): add benchmarking tool to docker image (#298 )	2023-05-09 13:19:31 +02:00
Sai Vinay G	926fd9a010	feat(router): Adding response schema for compat_generate (#292 )	2023-05-09 12:38:09 +02:00
Nicolas Patry	411b0d4e1f	chore(github): add templates (#264 )	2023-05-02 15:43:19 +02:00
Nicolas Patry	db2b4e0754	feat(router): new healthcheck that skips the queue (#244 ) Co-authored-by: OlivierDehaene <23298448+OlivierDehaene@users.noreply.github.com> Co-authored-by: OlivierDehaene <olivier@huggingface.co>	2023-04-26 20:23:54 +02:00
Nicolas Patry	c4fb09f2ae	feat(router): add tests to validation (#237 )	2023-04-26 16:14:40 +02:00
OlivierDehaene	8b182eb986	feat(router): add endpoint info to /info route (#228 )	2023-04-25 13:11:18 +02:00
OlivierDehaene	ebc74d5666	feat(router): use number of tokens in batch as input for dynamic batching (#226 ) Co-authored-by: Nick Hill <nickhill@us.ibm.com>	2023-04-24 17:59:00 +02:00
OlivierDehaene	343437c7b5	feat(router): add device and dtype info (#215 )	2023-04-21 15:36:29 +02:00
OlivierDehaene	709d8936f6	feat(router): drop requests when client closes the channel (#202 )	2023-04-20 11:07:40 +02:00
OlivierDehaene	2475aede61	feat(router): add info route (#196 ) close #125	2023-04-18 16:16:06 +02:00
OlivierDehaene	9987960062	feat(router): make router input validation optional (#164 )	2023-04-09 20:22:27 +02:00
OlivierDehaene	7dec65a244	fix(router): use buckets for metrics histograms (#163 )	2023-04-09 20:13:28 +02:00
OlivierDehaene	d503e8f09d	feat: aws sagemaker compatible image (#147 ) The only difference is that now it pushes to registry.internal.huggingface.tech/api-inference/community/text-generation-inference/sagemaker:... instead of registry.internal.huggingface.tech/api-inference/community/text-generation-inference:sagemaker-... --------- Co-authored-by: Philipp Schmid <32632186+philschmid@users.noreply.github.com>	2023-03-29 21:38:30 +02:00
OlivierDehaene	55bd4fed7d	feat(router): add best_of parameter (#117 )	2023-03-09 15:30:54 +01:00
OlivierDehaene	e8bfe199ba	feat(router): support left truncation (#115 ) closes #111	2023-03-09 13:10:30 +01:00
OlivierDehaene	1a2d68250a	feat: support typical sampling (#114 ) closes #112	2023-03-09 11:33:57 +01:00
OlivierDehaene	3fef90d50f	feat(clients): Python client (#103 )	2023-03-07 18:52:22 +01:00
OlivierDehaene	9b8ea6a6c7	feat(server): add logits watermark (#90 )	2023-03-02 12:30:41 +01:00
OlivierDehaene	f874c47831	feat(router): add api-inference headers (#91 )	2023-03-02 11:41:51 +01:00
OlivierDehaene	4e685d907e	feat(router): ask hf.co for pipelinetag to decide on compat_return_full_text (#89 )	2023-02-28 10:19:32 +01:00
OlivierDehaene	21340f24ba	feat(router): add legacy route for api-inference support (#88 )	2023-02-27 14:56:58 +01:00
OlivierDehaene	0ac184ce77	feat(server): add special token bool (#85 )	2023-02-24 15:55:57 +01:00
OlivierDehaene	6796d38c6d	feat(router): add cors allow origin options (#73 )	2023-02-17 18:22:00 +01:00
OlivierDehaene	439fcaf810	feat(router): add prometheus metrics scrape endpoint (#71 )	2023-02-16 17:18:53 +01:00
OlivierDehaene	5437d49beb	feat(router): add max_total_tokens and empty_input validation (#68 ) closes #65	2023-02-15 21:56:59 +01:00
OlivierDehaene	9af454142a	feat: add distributed tracing (#62 )	2023-02-13 13:02:45 +01:00
Yannic Kilcher	e520d5b349	fixed SSE naming (#61 ) https://en.wikipedia.org/wiki/Server-sent_events	2023-02-08 22:30:11 +01:00
OlivierDehaene	20c3c5940c	feat(router): refactor API and add openAPI schemas (#53 )	2023-02-03 12:43:37 +01:00
OlivierDehaene	b1482d9048	breaking(router): modify /generate API to only return generated text (#50 ) @njhill, @yk FYI generated_text was concatenated to the user prompt for legacy reason. We want to remove this behaviour as we don't think it is useful and even detrimonial to usability. We also remove the unused Vec.	2023-02-02 15:02:04 +01:00
OlivierDehaene	313194f6d7	feat(server): support repetition penalty (#47 )	2023-02-01 15:58:42 +01:00
OlivierDehaene	017a2a8c2f	feat: Add token streaming using ServerSideEvents support (#41 )	2023-01-31 17:04:00 +01:00
OlivierDehaene	4f9ac67cfa	Revert "feat: Add token streaming using ServerSideEvents support" (#40 ) Reverts huggingface/text-generation-inference#36	2023-01-31 14:21:51 +01:00
OlivierDehaene	7fbfbb0dc5	feat: Add token streaming using ServerSideEvents support (#36 ) Add token streaming using ServerSideEvents (SSE). The signature of the SSE events is: ```rust struct Details { finish_reason: String, generated_tokens: u32, seed: Option<u64>, } struct StreamResponse { token: Token, generated_text: Option<String>, details: Option<Details>, } struct ErrorResponse { error: String, } ```	2023-01-31 11:49:43 +01:00
OlivierDehaene	cd298bc5e5	feat: Support sampling seeding (#37 ) Co-authored-by: Yannic Kilcher <yk@users.noreply.github.com>	2023-01-30 15:36:16 +01:00
OlivierDehaene	5c01e2544c	fix(router): fix api-inference deployment (#31 )	2023-01-23 17:42:14 +01:00
OlivierDehaene	f9d0ec376a	feat(docker): Make the image compatible with api-inference (#29 )	2023-01-23 17:11:27 +01:00
OlivierDehaene	32a253063d	feat: Return logprobs (#8 )	2022-12-15 17:03:56 +01:00
OlivierDehaene	718096f695	feat: Support stop sequences (#7 )	2022-12-12 18:25:22 +01:00
Nick Hill	31d76e238d	fix(batching): Avoid theoretical hang in batcher loop (#5 ) - Avoid theoretical hang in batcher loop - Avoid a couple of clones in the router generate method - Keep attention mask tensors as integers - Remove num_heads attribute Co-authored-by: OlivierDehaene <Olivier.dehaene@gmail.com>	2022-12-05 10:10:59 +01:00
OlivierDehaene	c5665f5c8b	feat(server): Support generic AutoModelForCausalLM	2022-11-04 14:22:47 +01:00

1 2

62 Commits