Add streaming to `profiling.sampling`
未關閉
還沒有人認領這個 Issue。
stdlib
topic-profiling
type-feature
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Feature or enhancement
Proposal:
Right now, profiling.sampling has roughly two modes: live with the TUI, and snapshot-at-the-end (except binary? but see the note below.) There's nothing that streams the data continously as it comes. This would be ideal for long-running headless profiling.
This one is less defined than #145411, so there are more open questions:
- Should it stream raw or agreggate data? What should be the window?
- What should be the format? Unfortunately, from what I checked the current binary format is not really well-suited for streaming, as it saves the dictionaries only on finalize.
- What should be the transport layers for streaming?
- What should be the types of messages?
- What should be the configuration flags?
- How the backpressure should be handled? Should it drop the oldest? All the oldest?
My hunch is:
- support both raw and aggregate and assume that aggregate is just a different message type.
- start with something simple as JSONL
- I have mixed feelings about the transport layer. The stdout sounds great on paper but it's a mixed-use channel right now and there's a question around blocking by slower consumers. Maybe for debugging? Unix socket is good but not perfectly portable
- as for backpressure, I would just drop the oldest by default, but only the
raw,aggandheartbeatmessages.
For starters:
--stream unix:/tmp/tachyon-stream.sockor--stream file:/tmp/tachyon-stream.sock(later:--stream tcp:127.0.0.1:1234)--stream-types meta|str_def|frame_def|raw|agg:5s|loss|heartbeat|end|error(comma separated; note theagg:5shere; we it could be extended toheartbeat:1setc. in the future)--stream-format jsonl
Then, in the next phases we could think of --stream-drop-policy etc.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
先閱讀 profiling.sampling 的實作,以及 issue #145411 中的相關討論。在縮小設計範圍之前,檢視提議的串流類型、JSONL 格式、傳輸方式和 backpressure 相關問題。確定已達成共識的串流範圍,以及支援長時間執行 headless profiling 的設定後,即可完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- performance
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 冷清
- 描述清晰度
- 需要釐清
- 新手友好度
- 25/100