python / python/cpython

asyncio: SelectorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF

Aperta
#156,344 4 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

stdlib topic-asyncio type-bug
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Bug description

The selector-side twin of #156333. BaseSelectorEventLoop also wakes itself through a self-socket created by socket.socketpair(). On Windows this is a loopback TCP pair, and if the connection reaches a clean EOF while the loop runs, _read_from_self returns but the reader is still registered on the dead fd — select() reports it readable immediately, forever, pinning one core at 100% with nothing logged.

Lib/asyncio/selector_events.py:

def _read_from_self(self):
    while True:
        try:
            data = self._ssock.recv(4096)
            if not data:
                break
            self._process_self_data(data)
        except InterruptedError:
            continue
        except BlockingIOError:
            break

At EOF, recv() returns b'', the loop breaks, and the callback returns. But the fd is still registered via _add_reader(self._ssock.fileno(), self._read_from_self) from _make_self_pipe(). A closed-for-read socket is permanently "readable", so every select() iteration re-fires the callback — a tight spin. The loop is otherwise idle and never recovers.

The proactor half of this bug is fixed by #156343 (rebuild the pair on the empty result). The same OS teardown of the loopback pair that triggers it triggers this on any WindowsSelectorEventLoopPolicy process.

Reproducer (deterministic, Windows)
import asyncio, socket, time

calls = 0
orig = asyncio.selector_events.BaseSelectorEventLoop._read_from_self
def counting(self):
    global calls
    calls += 1
    return orig(self)
asyncio.selector_events.BaseSelectorEventLoop._read_from_self = counting

async def main():
    loop = asyncio.get_running_loop()
    await asyncio.sleep(0.1)
    loop._csock.shutdown(socket.SHUT_WR)   # graceful half-close = clean EOF
    global calls
    calls = 0
    start = time.process_time()
    await asyncio.sleep(3)
    print(f"_read_from_self calls during 3s idle: {calls}")
    print(f"CPU consumed while sleeping 3s: {time.process_time() - start:.2f}s")

loop = asyncio.SelectorEventLoop()
asyncio.set_event_loop(loop)
loop.run_until_complete(main())

Measured on Windows 11, CPython main (3.16.0a0, self-built): 582,692 callback invocations, 1.23s CPU, during a 3-second idle sleep. Expected ~0.

Note the instrumentation must patch the class before the loop is created — _make_self_pipe() captures the bound method at registration time, so patching an instance afterwards does not intercept the already-armed reader (which is also why the spin is invisible to profilers that hook late).

Environment
  • Windows 11, self-built main (3.16.0a0)
  • The proactor variant (#156333) reproduces on installed 3.12.10 / 3.13.12; the _read_from_self code path shown above is unchanged on main
Notes on scope
  • Unix loops use an AF_UNIX pair, which the OS does not tear down this way, so the organic trigger is Windows-specific; a deterministic test using shutdown(SHUT_WR) works on any platform, though.
  • Related, on the proactor side: the same teardown can also surface as ConnectionResetError from the pending recv instead of a clean EOF; that path is currently unhandled too (noted in #156343).
Suggested direction

On recv() == b'', rebuild the self-pipe the way #156343 does for the proactor loop: allocate the new pair first, move any signal wakeup fd registration (set_wakeup_fd() returns the previous fd, so it can be probed and moved only when it names the old socket — Unix loops with signal handlers), remove the old reader, close the old sockets, and register the reader on the new fd.

Linked PRs
  • gh-156345

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia in Lib/asyncio/selector_events.py con _read_from_self e _make_self_pipe, quindi confronta la correzione del proactor in #156343 e il lavoro collegato in #156345. Riproduci il caso di EOF del self-pipe su Windows e verifica che un selector loop inattivo non invochi più ripetutamente il reader, mentre la registrazione del signal wakeup rimane corretta.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend, operating-systems
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Specificata chiaramente
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.