Lightning-AI / Lightning-AI/lightning-thunder
CI: Re-Enable torchrun call in Zero to Thunder notebook
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 121
- PR merge metrics
- No merged PRs in 30d
Description
Because the CI runs into flakiness problems with distributed, I am disabling the call to torchrun in #452 , it would be neat to re-enable once we know what is going on.
Based on analysis of the build-logs, the problem seems to be connected to lambda-server1 , but it could also be some other but correlated thing. (This is off the 186 runs I got from the Azure API this morning.)

cc @borda
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the Zero to Thunder notebook change in #452 and the available CI build logs, including the 186-run Azure analysis. Investigate whether lambda-server1 or another correlated factor causes the distributed flakiness; this is done when the torchrun call can be re-enabled and the CI runs reliably.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- ci-cd, distributed-systems, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100