Use dual stacks to separate control and data
まだ誰も着手していません。
- 主要言語
- Python
- スター
- 77.2k
- フォーク
- 35.9k
- PR マージ指標
- PR 指標を取得中
説明
Currently, in CPython, the stack is implemented as a linked list of frames. These frames are mostly allocated in large chunks, but may not be contiguous as generator frames are allocated as part of the generator.
Each frame's size depends on its code object, which makes scanning the scan slow and requires dynamic memory allocation to account for varying stack sizes, making it hard for external profilers, like Tachyon, to snapshot the stack without relying on undocumented assumptions about the VM.
Tachyon is advertised as zero-overhead, but it isn't. It requires extra code on the fast path of yields and returns and couples the VM and profiler in undocumented and hard to maintain ways preventing improvements in the VM e.g. https://github.com/python/cpython/pull/148681
By splitting the stack into two parts, a control stack and an object stack, we can make the control stack simpler, smaller and more regular, which would:
- make it easier and faster for Tachyon to copy and parse, and keep the object stack layout flexible, and
- give the VM freedom to arrange the object stack however is best for performance.
Performance impact
Having two stacks, means two stack pointers, using a register in the interpreter and JIT.
However, having a control stack of fixed-sized control frames, would simplify recursion depth checking and other bookkeeping tasks.
I expect the additional costs and the savings to largely cancel out.
Primarily, this is about decoupling profilers from the VM design, not performance, so a small initial slowdown would be acceptable.
Having the object stack being purely composed of object pointers, and no control, may allow some additional optimizations across calls, slightly reduced stack memory use, and slightly faster stack scanning for the GC, so this might eventually give a small performance boost.
Control frames
Profiler view
To a profiler, or other out-of-process tool, a control frame will look like this:
typedef struct {
PyObject *executable; // The code object, or non-Python callable, for this frame.
char padding[CONTROL_FRAME_SIZE - sizeof(void *)];
} _PyControlFrame;
If callable is a code object, then additional information is available:
typedef struct {
PyCodeObject *code; // The code for this frame
_Py_CODEUNIT *instr_ptr; /* Instruction currently executing (may be approximate) */
char padding2[CONTROL_FRAME_SIZE - 2 * sizeof(void *)];
} _PyPythonControlFrame;
Implementation
The control frame contains all the data that isn't object pointers from the current interpreter frame, plus the pointer to the code object:
typedef struct {
PyObject *executable; /* Borrowed reference */
_Py_CODEUNIT *instr_ptr;
_PyInterpreterDataFrame *framepointer;
_PyStackRef *stackpointer;
uint8_t owner;
uint8_t visited; /* For GC */
uint16_t return_offset; /* Only relevant during a function call */
/* 32 bits used for tlbc_index in FT build */
} _PyControlFrameInternal;
typedef struct {
_PyStackRef f_funcobj; /* Deferred or strong reference. */
_PyStackRef f_globals; /* Borrowed reference. */
_PyStackRef f_builtins; /* Borrowed reference. */
_PyStackRef f_locals; /* Strong reference, may be NULL. */
_PyStackRef frame_obj; /* Strong reference, may be NULL. */
/* Locals and stack */
_PyStackRef localsplus[1];
} _PyInterpreterDataFrame;
Debug info and guarantees for profilers
Out of process debuggers and profilers require information to traverse internal data structures. The VM will provide this information:
uint32_t control_frame_offsetoffset of the control frame pointer in the thread stateuint32_t control_base_offsetoffset of the pointer to the base of the stack in the thread stateuint32_t control_frame_sizethe size (in bytes) of a control frame ==CONTROL_FRAME_SIZEabove.
Any changes to the base pointer changes will be protected by a memory fence, so other processors will see the change.
The control frame pointer will not be synchronized, so may appear out-of-date to other processors.
Prior discussion focused on implementation: https://github.com/faster-cpython/ideas/issues/675
https://github.com/python/cpython/issues/115946 explains why this helps profilers.
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
まず、faster-cpython/ideas#675 の以前の実装に関する議論と、CPython#115946 の profiler のコンテキストを読んでください。Issue では実装ファイルやテストは特定されていません。完了条件は、control stack と object stack の分離を定義して実装し、記載されている profiler のトラバーサル保証を維持し、パフォーマンス上のトレードオフを評価することです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- compilers, devtools
- issue の種類
- リファクタリング
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 30/100