NVIDIA / NVIDIA/stdexec

replace the lock in `exec::__nest_rcvr::__complete` with a CAS loop

Open
#1,197 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement P1
Dominant language
C++
Stars
2.4k
Forks
270
Avg merge
3d 6h
Merged PRs (30d)
39

Description

async_scope::spawn is much slower than start_detached, and I suspect the issue is the locking and unlocking going on in exec::__next_rcvr::__complete. There is similar logic for walking a linked list and notifying each element in the implementation of split, but there it is done more efficiently with a compare-and-swap loop. Maybe the two can share logic.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with async_scope::spawn and start_detached, then inspect the linked-list notification logic in split. Locate the locking and unlocking in exec::__next_rcvr::__complete (or the __nest_rcvr::__complete named in the title) and compare it with split’s compare-and-swap loop. Done means the completion path uses the CAS-based approach without changing notification behavior or async_scope correctness.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance
Issue type
Refactor
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.