Optimize reference tracking and eliminate branching during returns and yields
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
There are a few optimizations we can make to returns and yields, but they are somewhat related so I'm grouping them into a single issue.
During a return, the VM needs to make the returned reference "heap safe" (convert any borrowed references to strong references) and then pop and clear the frame, which can involve quite a lot of decrefs.
During a yield, the VM needs to make the returned reference heap safe and then pop the frame, but not clear it.
We want to minimize the amount of refcounting operations that we do.
Before we can do much else, we should split RETURN_VALUE and YIELD_VALUE into micro-ops:
macro(RETURN_VALUE) = _MAKE_HEAP_SAFE + _RETURN_VALUE
macro(YIELD_VALUE) = _MAKE_HEAP_SAFE + _YIELD_VALUE
so that we can optimize the uops independently
Eliminate making heap safe if we know that reference already is
If TOS is a strong reference _MAKE_HEAP_SAFE -> _NOP
Avoid the branch in _RETURN_VALUE by splitting into _RETURN_VALUE_GEN and _RETURN_VALUE_FUNC
To avoid using up opcodes, we'll need to do this in the JIT, not the interpreter.
Track borrows/immortals to optimize clearing frames
E.g. if a frame has four local variables, a, b, c, d but we can tell that a and d are immortal, or borrowed, then instead of looping over all the variables, we could emit code just to decref b and c.
This might end up bloating the code, as we would need to inline the decrefs, but it could be quite a lot faster.
Linked PRs
- gh-144414
- gh-146320
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the RETURN_VALUE and YIELD_VALUE macros, then trace _MAKE_HEAP_SAFE, _RETURN_VALUE, and _YIELD_VALUE through the interpreter and JIT. Compare the linked work gh-144414 and gh-146320 before deciding scope; done would require independently optimized return and yield paths and reduced unnecessary refcounting without regressing behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- compilers, performance
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100