python / python/cpython

asyncio.Barrier: cancelling a waiting task mid-round can cause a later arrival to reuse its index

未關閉
#155,233 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

stdlib topic-asyncio type-bug
主要語言
Python
星號
77.2k
分支
35.9k
PR 合併指標
PR 指標待擷取

描述

Bug description:

asyncio.Barrier.wait() is documented to return "a unique and individual index number from 0 to 'parties-1'" for each task that passes the barrier together. If a task's wait() call is cancelled while the barrier is still filling (i.e. before enough parties have arrived), a task that arrives afterward can be assigned the same index as another task that is already waiting in that same round — so two tasks in the same successful release get the same index, and another index in the range is never handed out at all.

Reproducer
import asyncio

async def main():
    b = asyncio.Barrier(3)
    results = []

    async def party(name):
        try:
            idx = await b.wait()
            results.append((name, idx))
        except asyncio.CancelledError:
            results.append((name, "cancelled"))
            raise

    t1 = asyncio.create_task(party("A"))
    t2 = asyncio.create_task(party("B"))
    await asyncio.sleep(0)
    await asyncio.sleep(0)

    t1.cancel()          # A leaves while the barrier is still filling
    try:
        await t1
    except asyncio.CancelledError:
        pass
    await asyncio.sleep(0)

    t3 = asyncio.create_task(party("C"))
    t4 = asyncio.create_task(party("D"))   # completes the round of 3
    await asyncio.gather(t2, t3, t4)

    print(results)

asyncio.run(main())

Output on current main:

[('A', 'cancelled'), ('D', 2), ('B', 1), ('C', 1)]

B and C both get index 1; index 0 is never handed out to any of the three tasks that actually pass the barrier together (B, C, D).

Root cause

In Lib/asyncio/locks.py, Barrier.wait() assigns index = self._count eagerly, at arrival, before the round's final membership is known:

index = self._count
self._count += 1
if index + 1 == self._parties:
    await self._release()
else:
    await self._wait()
return index

self._count is decremented again when a party leaves early (cancellation), in the finally block. Since index is derived from a counter that can both increase (new arrivals) and decrease (early departures) within the same round, a later arrival can end up with the exact self._count value an earlier, still-waiting task already captured as its own index.

This is specific to the asyncio implementation — threading.Barrier doesn't need to handle a party "un-arriving" mid-wait the way a cancellable asyncio.Task can, so the sync version doesn't have an analogous bug.

I don't see an existing report for this (checked issues/PRs mentioning asyncio.Barrier).

A working fix

Indices can't be assigned correctly at arrival time when membership can still change; they need to be finalized once the round's exact release cohort is known (at release time), assigned once from the current arrival order. I have a fix + regression test ready and will open a PR against this issue.

CPython versions tested on:

3.16 (main)

Operating systems tested on:

Linux (WSL), should be platform-independent (no OS-specific code involved)

Linked PRs
  • gh-155235

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 Lib/asyncio/locks.py 中的 Barrier.wait 開始,執行 issue 中描述的 reproducer,以觀察取消後重複的 index。檢視連結的 PR gh-155235 及其 regression test。完成的定義是,取消情境會釋放一個 round,且每個通過的 task 都具有唯一的 index。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
backend
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
描述清楚
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。