991 B
991 B
Using TGI with Intel GPUs
TGI optimized models are supported on Intel Data Center GPU Max1100, Max1550, the recommended usage is through Docker.
On a server powered by Intel GPUs, TGI can be launched with the following command:
model=teknium/OpenHermes-2.5-Mistral-7B
volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run
docker run --rm --privileged --cap-add=sys_nice \
--device=/dev/dri \
--ipc=host --shm-size 1g --net host -v $volume:/data \
ghcr.io/huggingface/text-generation-inference:2.2.0-intel \
--model-id $model --cuda-graphs 0
The launched TGI server can then be queried from clients, make sure to check out the Consuming TGI guide.