python / python/cpython

Allocate JIT memory in large chunks near the executable

Open
#148,822 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

interpreter-core topic-JIT type-feature
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

By allocating JIT memory in large chunks we can reduce the overhead of compilation and possibly speed up the compiled code a bit.

Currently, when we need to allocate memory for the JIT, we ask the OS for a sufficiently large chunk of virtual memory. This is simple and reasonably efficient. It does have a few flaws:

  1. We need to create DWARF debug info and register it for each trace
  2. We need to make a syscall to get memory for each trace
  3. We have no control over the location of the jitted code, meaning that calls in the executable may need trampolines

By allocating large chunks we can reduce the overhead of (1) and (2) to once per-chunk, not per trace.
By allocating large chunks we can also afford the additional overhead of requesting memory near to the executable.

How it would work:

  • When we need memory for jitted code, we request it from our special allocator.
  • When the allocator needs memory, it requests it from the OS, making several requests for it near the executable before accepting an location
  • The allocator itself will be a standard obmalloc/jemalloc style block allocator
Size classes and fragmentation:

All jitted code will need to page aligned, so blocks will need to a multiple of the page size.
With 4 size classes per power of 2 increase in size (like jemalloc) and assuming 2M chunks we get relatively little internal fragmentation, but potential quite a lot of external fragmentation as there would be around 20 size classes.
The external fragmentation can be mitigated by allowing the OS to lazily allocate the pages on demand.

API

The _PyObject_VirtualAlloc function will need extending (or a new function added) to allow the desired address to be passed, so the allocator can get blocks near the executable.

The jit_alloc function's API will be unchanged.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing the JIT allocation path through _PyObject_VirtualAlloc and jit_alloc, then review the existing virtual-memory allocation behavior. The work is complete when JIT memory is obtained from large, page-aligned chunks near the executable, with per-chunk DWARF registration and the jit_alloc API unchanged.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, python
Domain
compilers, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.