NVIDIA / NVIDIA/cuda-python

cuda.core: GraphBuilder.join leaves forked builders in a state that segfaults at garbage collection if it raises midway

Offen
#2,776 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

triage
Vorherrschende Sprache
Cython
Sterne
3.4k
Forks
329
Ø Merge
1 T. 23 Std.
Gemergte PRs (30 T.)
116

Beschreibung

Summary

If GraphBuilder.join raises partway through, the forked builders it has not yet closed are left open with capturing streams. The test that hit this failed cleanly, but the process then segfaulted during gc.collect() in the init_cuda fixture teardown while those abandoned objects were destroyed. An exception inside join should not be able to crash the interpreter later.

What was observed

On PR #2750, a transient bug made the temporary ordering event in Stream.wait fail to be created. join calls root_bdr.stream.wait(builder.stream) and then builder.close() for each non-root builder; the wait raised on the first builder, so no forked builder was closed. Every GPU test job then crashed with:

Fatal Python error: Segmentation fault
Current thread ... (most recent call first):
  File ".../cuda_core/tests/conftest.py", line 192 in init_cuda

Line 192 is the gc.collect() in the fixture's finally. The first tests to fail were test_graph_complete_after_close_forked and test_graph_definition_raises_for_forked in tests/graph/test_graph_builder.py, both of which go through split and join.
https://github.com/NVIDIA/cuda-python/actions/runs/33911046695/job/101149369210

The event-creation bug is fixed in that PR, so the crash is no longer reachable through this route. The teardown fragility remains: any exception in join (or an interrupted split/join sequence) leaves builders in the same state.

Suggested direction

  • Make join exception-safe: on failure, close or otherwise neutralize the forked builders that were not joined, rather than leaving them mid-capture.
  • Make the forked-builder destructor tolerant of an abandoned capture, so destruction of a never-joined fork cannot dereference an invalid handle or end capture on a stream that is no longer valid. Identifying the exact dereference is part of this issue; the Python traceback stops at gc.collect().

Reproduction sketch

gb = Device().create_graph_builder().begin_building()
left, right = gb.split(2)
# force root_bdr.stream.wait(...) to raise inside join, e.g. by monkeypatching Stream.wait
with pytest.raises(Exception):
    GraphBuilder.join(left, right)
del left, right, gb
gc.collect()   # crashes today

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne mit GraphBuilder.join und den von tests/graph/test_graph_builder.py ausgeübten Bereinigungspfaden des Forked-Builders und untersuche anschließend den init_cuda-Teardown in cuda_core/tests/conftest.py in Zeile 192. Verwende die Reproduktionsskizze, damit Stream.wait eine Exception auslöst, führe die beiden benannten Graph-Builder-Tests aus und überprüfe, dass fehlgeschlagene Joins keine verwaisten Captures hinterlassen und gc.collect() keinen Segmentation Fault mehr verursacht.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
hpc
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
48/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.