python / python/cpython

`_Py_Dealloc()` being unaware of separate stacks can cause memory leaks

Abierto
#157,519 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

type-bug
Lenguaje dominante
Python
Estrellas
77.2k
Forks
35.9k
Métricas de merge de PR
Métricas de PR pendientes

Descripción

Bug report

Bug description:

To avoid unbounded deallocations (which would cause stack overflows) _Py_Dealloc() calls _Py_RecursionLimit_GetMargin() to check how much stack space is available, and if there's not enough stack space then it will put the object on a queue to free later. The problem is that when _Py_RecursionLimit_GetMargin() is called under a separate userspace stack it will calculate something very negative, so _Py_Dealloc() will never free anything and just keep adding objects to the trash queue.

This came up in practice when calling Python on a Julia task through PythonCall.jl, but here's a pure Python MWE that simulates the situation by running a workload under different stack situations:

import ctypes
import gc
import sys
import tracemalloc


api = ctypes.pythonapi
api.PyThreadState_Get.restype = ctypes.c_void_p
api.PyUnstable_ThreadState_SetStackProtection.argtypes = [
    ctypes.c_void_p, ctypes.c_void_p, ctypes.c_size_t]
api.PyUnstable_ThreadState_ResetStackProtection.argtypes = [ctypes.c_void_p]

# Real bounds of the current stack stack
libc = ctypes.CDLL(None, use_errno=True)
libc.pthread_self.restype = ctypes.c_void_p
attr = ctypes.create_string_buffer(1024)
assert libc.pthread_getattr_np(ctypes.c_void_p(libc.pthread_self()), attr) == 0
stack_addr = ctypes.c_void_p()
stack_size = ctypes.c_size_t()
assert libc.pthread_attr_getstack(attr, ctypes.byref(stack_addr), ctypes.byref(stack_size)) == 0
libc.pthread_attr_destroy(attr)

# Pretend the stack is 1 GiB above where it really is
fake_size = 1 << 20
fake_start = stack_addr.value + stack_size.value + (1 << 30)

deleted = 0
N = 1000

# Dummy class that counts how many times the destructor was called
class Foo:
    def __init__(self):
        # Allocate some memory
        self.payload = bytes(100_000)

    def __del__(self):
        global deleted
        deleted += 1

def churn():
    for _ in range(N):
        Foo() # refcount hits zero immediately


def report(label):
    gc.collect()
    current, _ = tracemalloc.get_traced_memory()
    print(f"{label:<28} __del__ calls: {deleted:5d}/{N}   "
          f"traced memory: {current / 2**20:7.1f} MiB")

# Normal case
tracemalloc.start()
churn()
report("real stack limits")

# Simulate running under a different stack
tstate = api.PyThreadState_Get()
assert api.PyUnstable_ThreadState_SetStackProtection(
    tstate, fake_start, fake_size) == 0
deleted = 0
churn()
report("stack limits far away")

# Go back to the original stack and dealloc something to trigger cleanup of the
# delete_later list.
api.PyUnstable_ThreadState_ResetStackProtection(tstate)
trigger = []
del trigger
report("limits restored")

On 3.14.2 this prints out:

real stack limits            __del__ calls:  1000/1000   traced memory:     0.0 MiB
stack limits far away        __del__ calls:     0/1000   traced memory:    95.5 MiB
limits restored              __del__ calls:  1000/1000   traced memory:     0.0 MiB

i.e. in the case of a userspace stack Foo's destructor is never called and its memory is never freed. This is related to the new stack overflow detection: https://github.com/python/cpython/issues/139653
On 3.14.1 I think the script would have just aborted because the detection did not support userspace stacks at all: https://github.com/python/cpython/pull/141944
This looks like the same issue: https://github.com/python/cpython/issues/144165
It also appeared in ray: https://github.com/ray-project/ray/issues/63290#issuecomment-4980526793

CPython versions tested on:

3.14

Operating systems tested on:

Linux

Linked PRs
  • gh-157520

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Revisa el PR vinculado gh-157520 junto con las rutas _Py_Dealloc y _Py_RecursionLimit_GetMargin descritas en el informe. Ejecuta el MWE de Python proporcionado en CPython 3.14 con límites de pila reales y simulados; se considera terminado cuando los objetos se desasignan y la memoria trazada se libera bajo stacks de userspace separados.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
operating-systems, performance
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.