cuda.core: GraphBuilder.join leaves forked builders in a state that segfaults at garbage collection if it raises midway
Nessuno ha ancora preso questa issue.
- Lingua principale
- Cython
- Stelle
- 3.4k
- Fork
- 329
- Merge medio
- 1g 23h
- PR unite (30g)
- 116
Descrizione
Summary
If GraphBuilder.join raises partway through, the forked builders it has not yet closed are left open with capturing streams. The test that hit this failed cleanly, but the process then segfaulted during gc.collect() in the init_cuda fixture teardown while those abandoned objects were destroyed. An exception inside join should not be able to crash the interpreter later.
What was observed
On PR #2750, a transient bug made the temporary ordering event in Stream.wait fail to be created. join calls root_bdr.stream.wait(builder.stream) and then builder.close() for each non-root builder; the wait raised on the first builder, so no forked builder was closed. Every GPU test job then crashed with:
Fatal Python error: Segmentation fault
Current thread ... (most recent call first):
File ".../cuda_core/tests/conftest.py", line 192 in init_cuda
Line 192 is the gc.collect() in the fixture's finally. The first tests to fail were test_graph_complete_after_close_forked and test_graph_definition_raises_for_forked in tests/graph/test_graph_builder.py, both of which go through split and join.
https://github.com/NVIDIA/cuda-python/actions/runs/33911046695/job/101149369210
The event-creation bug is fixed in that PR, so the crash is no longer reachable through this route. The teardown fragility remains: any exception in join (or an interrupted split/join sequence) leaves builders in the same state.
Suggested direction
- Make
joinexception-safe: on failure, close or otherwise neutralize the forked builders that were not joined, rather than leaving them mid-capture. - Make the forked-builder destructor tolerant of an abandoned capture, so destruction of a never-joined fork cannot dereference an invalid handle or end capture on a stream that is no longer valid. Identifying the exact dereference is part of this issue; the Python traceback stops at
gc.collect().
Reproduction sketch
gb = Device().create_graph_builder().begin_building()
left, right = gb.split(2)
# force root_bdr.stream.wait(...) to raise inside join, e.g. by monkeypatching Stream.wait
with pytest.raises(Exception):
GraphBuilder.join(left, right)
del left, right, gb
gc.collect() # crashes today
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia da GraphBuilder.join e dai percorsi di pulizia del forked-builder esercitati da tests/graph/test_graph_builder.py, quindi esamina il teardown di init_cuda alla riga 192 di cuda_core/tests/conftest.py. Usa lo schema di riproduzione per fare in modo che Stream.wait sollevi un'eccezione, esegui i due test di graph-builder indicati e verifica che i join falliti non lascino capture abbandonati e che gc.collect() non provochi più un segmentation fault.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- hpc
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Attiva
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 48/100