python / python/cpython

`_Py_Dealloc()` being unaware of separate stacks can cause memory leaks

Aberta
#157,519 0 comentários 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

type-bug
Linguagem predominante
Python
Estrelas
77.2k
Forks
36k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

Bug report

Bug description:

To avoid unbounded deallocations (which would cause stack overflows) _Py_Dealloc() calls _Py_RecursionLimit_GetMargin() to check how much stack space is available, and if there's not enough stack space then it will put the object on a queue to free later. The problem is that when _Py_RecursionLimit_GetMargin() is called under a separate userspace stack it will calculate something very negative, so _Py_Dealloc() will never free anything and just keep adding objects to the trash queue.

This came up in practice when calling Python on a Julia task through PythonCall.jl, but here's a pure Python MWE that simulates the situation by running a workload under different stack situations:

import ctypes
import gc
import sys
import tracemalloc


api = ctypes.pythonapi
api.PyThreadState_Get.restype = ctypes.c_void_p
api.PyUnstable_ThreadState_SetStackProtection.argtypes = [
    ctypes.c_void_p, ctypes.c_void_p, ctypes.c_size_t]
api.PyUnstable_ThreadState_ResetStackProtection.argtypes = [ctypes.c_void_p]

# Real bounds of the current stack stack
libc = ctypes.CDLL(None, use_errno=True)
libc.pthread_self.restype = ctypes.c_void_p
attr = ctypes.create_string_buffer(1024)
assert libc.pthread_getattr_np(ctypes.c_void_p(libc.pthread_self()), attr) == 0
stack_addr = ctypes.c_void_p()
stack_size = ctypes.c_size_t()
assert libc.pthread_attr_getstack(attr, ctypes.byref(stack_addr), ctypes.byref(stack_size)) == 0
libc.pthread_attr_destroy(attr)

# Pretend the stack is 1 GiB above where it really is
fake_size = 1 << 20
fake_start = stack_addr.value + stack_size.value + (1 << 30)

deleted = 0
N = 1000

# Dummy class that counts how many times the destructor was called
class Foo:
    def __init__(self):
        # Allocate some memory
        self.payload = bytes(100_000)

    def __del__(self):
        global deleted
        deleted += 1

def churn():
    for _ in range(N):
        Foo() # refcount hits zero immediately


def report(label):
    gc.collect()
    current, _ = tracemalloc.get_traced_memory()
    print(f"{label:<28} __del__ calls: {deleted:5d}/{N}   "
          f"traced memory: {current / 2**20:7.1f} MiB")

# Normal case
tracemalloc.start()
churn()
report("real stack limits")

# Simulate running under a different stack
tstate = api.PyThreadState_Get()
assert api.PyUnstable_ThreadState_SetStackProtection(
    tstate, fake_start, fake_size) == 0
deleted = 0
churn()
report("stack limits far away")

# Go back to the original stack and dealloc something to trigger cleanup of the
# delete_later list.
api.PyUnstable_ThreadState_ResetStackProtection(tstate)
trigger = []
del trigger
report("limits restored")

On 3.14.2 this prints out:

real stack limits            __del__ calls:  1000/1000   traced memory:     0.0 MiB
stack limits far away        __del__ calls:     0/1000   traced memory:    95.5 MiB
limits restored              __del__ calls:  1000/1000   traced memory:     0.0 MiB

i.e. in the case of a userspace stack Foo's destructor is never called and its memory is never freed. This is related to the new stack overflow detection: https://github.com/python/cpython/issues/139653
On 3.14.1 I think the script would have just aborted because the detection did not support userspace stacks at all: https://github.com/python/cpython/pull/141944
This looks like the same issue: https://github.com/python/cpython/issues/144165
It also appeared in ray: https://github.com/ray-project/ray/issues/63290#issuecomment-4980526793

CPython versions tested on:

3.14

Operating systems tested on:

Linux

Linked PRs
  • gh-157520

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Revise o PR vinculado gh-157520 juntamente com os caminhos _Py_Dealloc e _Py_RecursionLimit_GetMargin descritos no relatório. Execute o MWE de Python fornecido no CPython 3.14 com limites de pilha reais e simulados; considera-se concluído quando os objetos são desalocados e a memória rastreada é liberada sob stacks de userspace separados.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
operating-systems, performance
Tipo de issue
Bug
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Estagnada
Clareza
Razoavelmente clara
Facilidade para iniciantes
25/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.