Nicolas Patry
49b4b33e80
feat(server): Update convert logic. ( #483 )
...
Should be more robust to shared tensors (ok when using
`from_pretrained). But forcing us to add new checks in our loading
code (since the chosen key to keep might be different from
`transformers`).
---------
Co-authored-by: Ubuntu <ubuntu@ip-172-31-41-161.ec2.internal>
2023-06-23 12:40:46 +02:00
OlivierDehaene
ece7ffa40a
feat(server): improve flash attention import errors ( #465 )
...
@lewtun, is this enough?
Closes #458
Closes #456
2023-06-19 09:53:45 +02:00
OlivierDehaene
62f91f78ac
feat(server): support vectorized warpers in flash causal lm ( #317 )
...
Co-authored-by: Joel Lamy-Poirier <joel.lamy-poirier@servicenow.com>
2023-05-26 12:30:27 +02:00
Nicolas Patry
b4aa87db58
fea(server): decrease convert RAM requirements ( #286 )
2023-05-05 17:57:02 +02:00
Nicolas Patry
690fc31757
fix(server): fix convert ( #284 )
2023-05-05 15:28:08 +02:00
Nicolas Patry
f08343d44d
fix(server): Removes the parallelism in file convertion (during download) ( #275 )
2023-05-04 15:22:54 +02:00
OlivierDehaene
3fef90d50f
feat(clients): Python client ( #103 )
2023-03-07 18:52:22 +01:00