NVIDIA / NVIDIA/cuda-python

cuda.core: VirtualMemoryResource grow path's Buffer._clear() re-triggers mr.deallocate() on an already-freed VA range

Abierto
#2,877 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

bug cuda.core
Lenguaje dominante
Cython
Estrellas
3.4k
Forks
329
Merge medio
1 d 23 h
PR fusionados (30 d)
116

Descripción

Summary

_grow_allocation_slow_path (cuda_core/cuda/core/_memory/_virtual_memory_resource.py:362) frees the old buffer's VA range manually (cuMemAddressFree(int(buf.handle), aligned_prev_size), line 457), then calls buf._clear() (line 460) with a comment saying this stops the old buffer's destructor from freeing it again.

Buffer._clear() (_buffer.pyx:262) does self._h_ptr.reset(). For a shared_ptr, reset() isn't a no-op invalidation — if buf holds the last reference, it runs the deleter immediately. This buffer's handle was created with mr=self (Buffer.from_handle(..., mr=self) in allocate()), so its deleter calls back into mr.deallocate() with the old, already-freed pointer and size. deallocate() then calls cuMemRetainAllocationHandle on a VA range that no longer exists, which fails with CUDA_ERROR_INVALID_VALUE. Under the error-handling policy landed in #2759, that failure is caught in the deleter and reported as a CUDAWarning instead of being silent; previously it was silently swallowed.

Evidence

Seen in CI, test_vmm_allocator_grow_allocation:

_virtual_memory_resource.py:292: CUDAWarning: mr.deallocate() failed during Buffer
destruction; the allocation may have leaked: CUDA_ERROR_INVALID_VALUE: This indicates
that one or more of the parameters passed to the API call is not within an acceptable
range of values.

Hypothesis, not yet confirmed by reproduction

Read from the code during review of #2759, not verified by running the test with added instrumentation.

Suggested fix direction

_clear()'s intent here seems to be "drop bookkeeping for a range this code already freed by hand," not "release the handle normally." Those need to be different operations: either give Buffer a way to drop its handle without invoking the resource's deallocate() (e.g. release the underlying pointer from the shared_ptr without running the deleter), or have the grow path swap in a handle whose deleter is already a no-op.

Refs: found during review of #2759.

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza con _grow_allocation_slow_path en cuda_core/cuda/core/_memory/_virtual_memory_resource.py y Buffer._clear() en _buffer.pyx. Ejecuta test_vmm_allocator_grow_allocation e inspecciona las rutas Buffer.from_handle() y deallocate() para verificar si al limpiar el buffer antiguo se invoca su Deleter. Está terminado cuando la ruta de crecimiento ya no intenta desasignar un rango de VA ya liberado ni emite la CUDAWarning indicada.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
backend
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Activo
Claridad
Bastante claro
Aptitud para principiantes
52/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.