deepspeedai / deepspeedai/DeepSpeed
[BUG] [0.8.1] INT8 model loading/inference issue
Open
@lekurile is already working on this.
Since May 12, 2023.
bug
inference
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
We conducted tests on OPT/GPTJ/GPT-Neox/BLOOM 7B INT8, these models are all producing garbage outputs on DeepSpeed 0.8.1
OPT model is NCCL communication issue
GPT-NeoX 20B is producing garbage
BLOOM-7B: shape '[1, 4, 32, 384]' is invalid for input of size 16384
How we tested?
We generated int8 checkpoints of the model and then loaded them back. Example of doing the same with DS inference test suite.
deepspeed --num_nodes 1 \
--num_gpus 8 \
inference-test.py \
--use_kernel \
--ds_inference \
--use_meta_tensor \
--name EleutherAI/gpt-neox-20b \
--checkpoint_path /tmp/ws/gpt-neox-20b/ \
--save_mp_checkpoint_path /tmp/ws/sharded-gpt-neox-20b/ \
--dtype int8
deepspeed --num_nodes 1 \
--num_gpus 8 \
inference-test.py \
--use_kernel \
--ds_inference \
--use_meta_tensor \
--name EleutherAI/gpt-neox-20b \
--checkpoint_path /tmp/ws/sharded-gpt-neox-20b/ \
--dtype int8
More info this.
https://github.com/microsoft/DeepSpeed/issues/2770
Creating a new issue to track the int8 checkpoint loading issue.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.