how to enable torch.compile when use Megatron-LM as training backend?
Open
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 312
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 2
Description
As described, torch.compile seems not compatible with megatron-lm, especially when I train the model with PP/TP/EP etc. How Can I Enable torch.compile for kernel optimization?
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are named. Start by reproducing the training setup with Megatron-LM and the PP/TP/EP combinations described, then investigate the interaction between torch.compile and the training backend. Done means establishing whether kernel optimization can be enabled and documenting or implementing a verified path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100