asyncio: SelectorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Bug description
The selector-side twin of #156333. BaseSelectorEventLoop also wakes itself through a self-socket created by socket.socketpair(). On Windows this is a loopback TCP pair, and if the connection reaches a clean EOF while the loop runs, _read_from_self returns but the reader is still registered on the dead fd — select() reports it readable immediately, forever, pinning one core at 100% with nothing logged.
Lib/asyncio/selector_events.py:
def _read_from_self(self):
while True:
try:
data = self._ssock.recv(4096)
if not data:
break
self._process_self_data(data)
except InterruptedError:
continue
except BlockingIOError:
break
At EOF, recv() returns b'', the loop breaks, and the callback returns. But the fd is still registered via _add_reader(self._ssock.fileno(), self._read_from_self) from _make_self_pipe(). A closed-for-read socket is permanently "readable", so every select() iteration re-fires the callback — a tight spin. The loop is otherwise idle and never recovers.
The proactor half of this bug is fixed by #156343 (rebuild the pair on the empty result). The same OS teardown of the loopback pair that triggers it triggers this on any WindowsSelectorEventLoopPolicy process.
Reproducer (deterministic, Windows)
import asyncio, socket, time
calls = 0
orig = asyncio.selector_events.BaseSelectorEventLoop._read_from_self
def counting(self):
global calls
calls += 1
return orig(self)
asyncio.selector_events.BaseSelectorEventLoop._read_from_self = counting
async def main():
loop = asyncio.get_running_loop()
await asyncio.sleep(0.1)
loop._csock.shutdown(socket.SHUT_WR) # graceful half-close = clean EOF
global calls
calls = 0
start = time.process_time()
await asyncio.sleep(3)
print(f"_read_from_self calls during 3s idle: {calls}")
print(f"CPU consumed while sleeping 3s: {time.process_time() - start:.2f}s")
loop = asyncio.SelectorEventLoop()
asyncio.set_event_loop(loop)
loop.run_until_complete(main())
Measured on Windows 11, CPython main (3.16.0a0, self-built): 582,692 callback invocations, 1.23s CPU, during a 3-second idle sleep. Expected ~0.
Note the instrumentation must patch the class before the loop is created — _make_self_pipe() captures the bound method at registration time, so patching an instance afterwards does not intercept the already-armed reader (which is also why the spin is invisible to profilers that hook late).
Environment
- Windows 11, self-built main (3.16.0a0)
- The proactor variant (#156333) reproduces on installed 3.12.10 / 3.13.12; the
_read_from_selfcode path shown above is unchanged on main
Notes on scope
- Unix loops use an AF_UNIX pair, which the OS does not tear down this way, so the organic trigger is Windows-specific; a deterministic test using
shutdown(SHUT_WR)works on any platform, though. - Related, on the proactor side: the same teardown can also surface as
ConnectionResetErrorfrom the pending recv instead of a clean EOF; that path is currently unhandled too (noted in #156343).
Suggested direction
On recv() == b'', rebuild the self-pipe the way #156343 does for the proactor loop: allocate the new pair first, move any signal wakeup fd registration (set_wakeup_fd() returns the previous fd, so it can be probed and moved only when it names the old socket — Unix loops with signal handlers), remove the old reader, close the old sockets, and register the reader on the new fd.
Linked PRs
- gh-156345
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 Lib/asyncio/selector_events.py 中的 _read_from_self 和 _make_self_pipe 開始,然後比較 #156343 中的 proactor 修正以及 #156345 中的相關工作。重現 Windows self-pipe EOF 情況,並驗證閒置的 selector loop 不再反覆呼叫 reader,同時 signal wakeup 註冊仍然正確。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- backend, operating-systems
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 停滯
- 描述清晰度
- 描述清楚
- 新手友好度
- 25/100