Lightning-AI / Lightning-AI/lightning-thunder

CI: Re-Enable torchrun call in Zero to Thunder notebook

Open
#465 0 comments 0 reactions 0 assignees View on GitHub
bug ci / tests
Dominant language
Python
Stars
1.5k
Forks
121
PR merge metrics
No merged PRs in 30d

Description

Because the CI runs into flakiness problems with distributed, I am disabling the call to torchrun in #452 , it would be neat to re-enable once we know what is going on.

Based on analysis of the build-logs, the problem seems to be connected to lambda-server1 , but it could also be some other but correlated thing. (This is off the 186 runs I got from the Azure API this morning.)

![image](https://github.com/Lightning-AI/lightning-thunder/assets/20787943/75fb310a-f3ee-45ae-9422-3b12452c80f3)

cc @borda

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the Zero to Thunder notebook change in #452 and the available CI build logs, including the 186-run Azure analysis. Investigate whether lambda-server1 or another correlated factor causes the distributed flakiness; this is done when the torchrun call can be re-enabled and the CI runs reliably.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ci-cd, distributed-systems, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.