cuda.core: GraphBuilder.join leaves forked builders in a state that segfaults at garbage collection if it raises midway
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Cython
- Estrellas
- 3.4k
- Forks
- 329
- Merge medio
- 1 d 21 h
- PR fusionados (30 d)
- 113
Descripción
Summary
If GraphBuilder.join raises partway through, the forked builders it has not yet closed are left open with capturing streams. The test that hit this failed cleanly, but the process then segfaulted during gc.collect() in the init_cuda fixture teardown while those abandoned objects were destroyed. An exception inside join should not be able to crash the interpreter later.
What was observed
On PR #2750, a transient bug made the temporary ordering event in Stream.wait fail to be created. join calls root_bdr.stream.wait(builder.stream) and then builder.close() for each non-root builder; the wait raised on the first builder, so no forked builder was closed. Every GPU test job then crashed with:
Fatal Python error: Segmentation fault
Current thread ... (most recent call first):
File ".../cuda_core/tests/conftest.py", line 192 in init_cuda
Line 192 is the gc.collect() in the fixture's finally. The first tests to fail were test_graph_complete_after_close_forked and test_graph_definition_raises_for_forked in tests/graph/test_graph_builder.py, both of which go through split and join.
https://github.com/NVIDIA/cuda-python/actions/runs/33911046695/job/101149369210
The event-creation bug is fixed in that PR, so the crash is no longer reachable through this route. The teardown fragility remains: any exception in join (or an interrupted split/join sequence) leaves builders in the same state.
Suggested direction
- Make
joinexception-safe: on failure, close or otherwise neutralize the forked builders that were not joined, rather than leaving them mid-capture. - Make the forked-builder destructor tolerant of an abandoned capture, so destruction of a never-joined fork cannot dereference an invalid handle or end capture on a stream that is no longer valid. Identifying the exact dereference is part of this issue; the Python traceback stops at
gc.collect().
Reproduction sketch
gb = Device().create_graph_builder().begin_building()
left, right = gb.split(2)
# force root_bdr.stream.wait(...) to raise inside join, e.g. by monkeypatching Stream.wait
with pytest.raises(Exception):
GraphBuilder.join(left, right)
del left, right, gb
gc.collect() # crashes today
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Comienza con GraphBuilder.join y las rutas de limpieza del forked-builder ejercitadas por tests/graph/test_graph_builder.py; después, inspecciona el teardown de init_cuda en la línea 192 de cuda_core/tests/conftest.py. Usa el esquema de reproducción para hacer que Stream.wait genere una excepción, ejecuta las dos pruebas de graph-builder indicadas y verifica que los joins fallidos no dejen captures abandonados y que gc.collect() ya no provoque un segmentation fault.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- hpc
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Activo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 48/100