Use dual stacks to separate control and data
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Currently, in CPython, the stack is implemented as a linked list of frames. These frames are mostly allocated in large chunks, but may not be contiguous as generator frames are allocated as part of the generator.
Each frame's size depends on its code object, which makes scanning the scan slow and requires dynamic memory allocation to account for varying stack sizes, making it hard for external profilers, like Tachyon, to snapshot the stack without relying on undocumented assumptions about the VM.
Tachyon is advertised as zero-overhead, but it isn't. It requires extra code on the fast path of yields and returns and couples the VM and profiler in undocumented and hard to maintain ways preventing improvements in the VM e.g. https://github.com/python/cpython/pull/148681
By splitting the stack into two parts, a control stack and an object stack, we can make the control stack simpler, smaller and more regular, which would:
- make it easier and faster for Tachyon to copy and parse, and keep the object stack layout flexible, and
- give the VM freedom to arrange the object stack however is best for performance.
Performance impact
Having two stacks, means two stack pointers, using a register in the interpreter and JIT.
However, having a control stack of fixed-sized control frames, would simplify recursion depth checking and other bookkeeping tasks.
I expect the additional costs and the savings to largely cancel out.
Primarily, this is about decoupling profilers from the VM design, not performance, so a small initial slowdown would be acceptable.
Having the object stack being purely composed of object pointers, and no control, may allow some additional optimizations across calls, slightly reduced stack memory use, and slightly faster stack scanning for the GC, so this might eventually give a small performance boost.
Control frames
Profiler view
To a profiler, or other out-of-process tool, a control frame will look like this:
typedef struct {
PyObject *executable; // The code object, or non-Python callable, for this frame.
char padding[CONTROL_FRAME_SIZE - sizeof(void *)];
} _PyControlFrame;
If callable is a code object, then additional information is available:
typedef struct {
PyCodeObject *code; // The code for this frame
_Py_CODEUNIT *instr_ptr; /* Instruction currently executing (may be approximate) */
char padding2[CONTROL_FRAME_SIZE - 2 * sizeof(void *)];
} _PyPythonControlFrame;
Implementation
The control frame contains all the data that isn't object pointers from the current interpreter frame, plus the pointer to the code object:
typedef struct {
PyObject *executable; /* Borrowed reference */
_Py_CODEUNIT *instr_ptr;
_PyInterpreterDataFrame *framepointer;
_PyStackRef *stackpointer;
uint8_t owner;
uint8_t visited; /* For GC */
uint16_t return_offset; /* Only relevant during a function call */
/* 32 bits used for tlbc_index in FT build */
} _PyControlFrameInternal;
typedef struct {
_PyStackRef f_funcobj; /* Deferred or strong reference. */
_PyStackRef f_globals; /* Borrowed reference. */
_PyStackRef f_builtins; /* Borrowed reference. */
_PyStackRef f_locals; /* Strong reference, may be NULL. */
_PyStackRef frame_obj; /* Strong reference, may be NULL. */
/* Locals and stack */
_PyStackRef localsplus[1];
} _PyInterpreterDataFrame;
Debug info and guarantees for profilers
Out of process debuggers and profilers require information to traverse internal data structures. The VM will provide this information:
uint32_t control_frame_offsetoffset of the control frame pointer in the thread stateuint32_t control_base_offsetoffset of the pointer to the base of the stack in the thread stateuint32_t control_frame_sizethe size (in bytes) of a control frame ==CONTROL_FRAME_SIZEabove.
Any changes to the base pointer changes will be protected by a memory fence, so other processors will see the change.
The control frame pointer will not be synchronized, so may appear out-of-date to other processors.
Prior discussion focused on implementation: https://github.com/faster-cpython/ideas/issues/675
https://github.com/python/cpython/issues/115946 explains why this helps profilers.
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
先閱讀 faster-cpython/ideas#675 中先前的實作討論,以及 CPython#115946 中 profiler 的背景脈絡。此 issue 未指定任何實作檔案或測試;完成表示定義並實作 control stack 與 object stack 的拆分、維持所述的 profiler traversal 保證,並評估效能取捨。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- compilers, devtools
- Issue 類型
- 重構
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 活躍
- 描述清晰度
- 基本清楚
- 新手友好度
- 30/100