deepspeedai / deepspeedai/DeepSpeedExamples
Why not just use zero3 inference to generate sequence in DeepSpeed Chat stage3 training?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 6.8k
- Forks
- 1.1k
- Avg merge
- 2d 16h
- Merged PRs (30d)
- 1
Description
DeepSpeed Chat use tensor parallelism via hybrid engine to generate sequence in stage3 training.
I wonder if just use zero3 inference for generation is ok? So that we don't need to transform model params between train and eval mode.
Any explanations about the design of stage3 training would be appreciated. Thanks.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the DeepSpeed Chat stage3 training design and the hybrid engine's tensor-parallel generation path. Compare that flow with the proposed ZeRO-3 inference approach and document whether it is supported, including the implications for parameter transformation between training and evaluation. Done means a maintainer-approved explanation or design decision; no files or tests are identified in the issue.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100