python / python/cpython

`TO_BOOL_INT` repeatedly misses for non-compact exact integers

Aberta
#155,486 4 comentários 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

interpreter-core performance type-bug
Linguagem predominante
Python
Estrelas
77.2k
Forks
35.9k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

Feature or enhancement

Proposal:

Versions

CPython 3.15.0b4, Ubuntu 24.04.4 LTS, gcc 13.3.0

Enhancement

I noticed that the specialization function for TO_BOOL and the guard used by TO_BOOL_INT accept different sets of integers.

_Py_Specialize_ToBool() selects TO_BOOL_INT for any exact integer:

if (PyLong_CheckExact(value)) {
    specialized_op = TO_BOOL_INT;
    goto success;
}

However, the TO_BOOL_INT macro uses _GUARD_TOS_INT, and that requires the integer to be compact:

op(_GUARD_TOS_INT, (value -- value)) {
    PyObject *value_o = PyStackRef_AsPyObjectBorrow(value);
    EXIT_IF(!_PyLong_CheckExactAndCompact(value_o));
}

As a result, a non-compact exact integer can cause TO_BOOL_INT to be selected and then immediately miss its guard.

This mismatch appears to have been introduced by GH-143759. Before the refactoring, TO_BOOL_INT had its own PyLong_CheckExact() guard. The refactoring replaced it with a macro using the shared _GUARD_TOS_INT, whose domain is narrower.

I see two alternative ways to fix this.

Option 1: narrow the specialization function

One option would be to change the integer check in _Py_Specialize_ToBool() so that it agrees with the existing guard:

if (_PyLong_CheckExactAndCompact(value)) {
    specialized_op = TO_BOOL_INT;
    goto success;
}

With this change, non-compact integers remain on the generic TO_BOOL path instead of repeatedly entering and missing TO_BOOL_INT.

My main concern with this option was whether the additional compactness check in the specialization function could regress the common case (compact int), so I benchmarked both compact and non-compact integers.

The benchmark target contained 100 TO_BOOL sites and was warmed up before each measurement:

start = time.perf_counter_ns()
for _ in range(10_000):
    target(value)
elapsed = time.perf_counter_ns() - start

Positive values mean that the patched build was faster:

Version Input Performance change
CPython 3.15 non-compact exact int +2.21% (95% CI: +1.64% to +2.86%)
CPython 3.15 compact exact int +0.31% (95% CI: −0.15% to +0.70%)
CPython main non-compact exact int +4.03% (95% CI: +1.43% to +7.10%)
CPython main compact exact int −0.23% (95% CI: −0.50% to +0.07%)

The non-compact case improved because it no longer repeatedly enters and misses TO_BOOL_INT. For compact integers, both confidence intervals include zero, so I did not find evidence that the stronger specialization check causes a regression.

Option 2: give TO_BOOL_INT an exact-int guard

The other option is to keep _Py_Specialize_ToBool() unchanged and add a guard that checks exact type without requiring compactness:

op(_GUARD_TOS_EXACT_INT, (value -- value)) {
    PyObject *value_o = PyStackRef_AsPyObjectBorrow(value);
    EXIT_IF(!PyLong_CheckExact(value_o));
}

The TO_BOOL_INT macro would then use the new guard:

macro(TO_BOOL_INT) =
    _GUARD_TOS_EXACT_INT +
    unused/1 +
    unused/2 +
    _TO_BOOL_INT +
    _POP_TOP_INT;

This would restore the specialization domain from before GH-143759.


I am not sure which of these two options is preferable, but the current mismatch seems worth fixing, so I am opening this issue to get feedback on which direction would be better.

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-155531
  • gh-155699
  • gh-155700

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece por Python/specialize.c em _Py_Specialize_ToBool() e por Python/bytecodes.c em TO_BOOL_INT e _GUARD_TOS_INT. Revise os PRs vinculados gh-155531, gh-155699 e gh-155700 e, em seguida, reproduza o benchmark de aquecimento indicado; está concluído quando os domínios de especialização e de guard concordarem sem misses repetidos.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
c, python
Domínio
compilers, performance
Tipo de issue
Funcionalidade
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Estagnada
Clareza
Razoavelmente clara
Facilidade para iniciantes
35/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.