deepseek-ai / deepseek-ai/DeepSeek-VL2

The idea of pipeline parallel strategy during Inference.

Open
#36 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.4k
Forks
1.8k
PR merge metrics
No merged PRs in 30d

Description

Hi, thanks for open-sourcing. Could you please share the idea of pipeline parallel strategy during Inference in deployment?

I find that the load unbalance of image preprocess and vision encoder cause the gpu bubble. And I noticed that you mentioned some fine-grained strategy to alleviate this in section 4.2. Hyperparameters and Infrastructures in the paper. Could you please share these ideas?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.