deepspeedai / deepspeedai/DeepSpeed
ZeroQuant quantization kernels and LKD
Open
@yaozhewei is already working on this.
Since Nov 4, 2022.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Hi,
I was trying out the compression library for ZeroQuant quantization (for GPT-J model). While I was able to compress the model, I didn't see any throughput/latency gain from the quantization during inference. I have a few questions regarding this:
- Do you guys have any guide to running inference on compressed models(especially ZeroQuant)? InferenceEngine only seems to support Mixture-of-Quantization but not ZeroQuant. I also tried int8 quantization without using compression module as shown in the code snippet below but end up getting
CUDA error: an illegal memory accesserror - Have you guys released the fused kernels for GeLU+Quantize and GeMM+dequantize proposed in the ZeroQuant paper yet?
- Any tentative release date for Layer-by-layer Knowledge Distillation?
- What's the motivation for multiplying quantized input by scale here? Wouldn't that dequantize inputs?
injection_policy={gptj_transformer:
module_inject.replace_policy.HFGPTJLayerPolicy}
model = deepspeed.init_inference(
model,
mp_size=world_size,
dtype=torch.int8,
quantization_setting=2,
replace_with_kernel_inject=True,
injection_policy=injection_policy,
)
Any help would be appreciated.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.