Add streaming to `profiling.sampling`
Nessuno ha ancora preso questa issue.
- Lingua principale
- Python
- Stelle
- 77.2k
- Fork
- 35.9k
- Metriche di merge delle PR
- Metriche PR in attesa
Descrizione
Feature or enhancement
Proposal:
Right now, profiling.sampling has roughly two modes: live with the TUI, and snapshot-at-the-end (except binary? but see the note below.) There's nothing that streams the data continously as it comes. This would be ideal for long-running headless profiling.
This one is less defined than #145411, so there are more open questions:
- Should it stream raw or agreggate data? What should be the window?
- What should be the format? Unfortunately, from what I checked the current binary format is not really well-suited for streaming, as it saves the dictionaries only on finalize.
- What should be the transport layers for streaming?
- What should be the types of messages?
- What should be the configuration flags?
- How the backpressure should be handled? Should it drop the oldest? All the oldest?
My hunch is:
- support both raw and aggregate and assume that aggregate is just a different message type.
- start with something simple as JSONL
- I have mixed feelings about the transport layer. The stdout sounds great on paper but it's a mixed-use channel right now and there's a question around blocking by slower consumers. Maybe for debugging? Unix socket is good but not perfectly portable
- as for backpressure, I would just drop the oldest by default, but only the
raw,aggandheartbeatmessages.
For starters:
--stream unix:/tmp/tachyon-stream.sockor--stream file:/tmp/tachyon-stream.sock(later:--stream tcp:127.0.0.1:1234)--stream-types meta|str_def|frame_def|raw|agg:5s|loss|heartbeat|end|error(comma separated; note theagg:5shere; we it could be extended toheartbeat:1setc. in the future)--stream-format jsonl
Then, in the next phases we could think of --stream-drop-policy etc.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Direzione di ricerca
Inizia leggendo l’implementazione di profiling.sampling e la discussione correlata nell’issue #145411. Esamina i tipi di stream proposti, il formato JSONL, i transports e le questioni relative alla backpressure prima di restringere il design. Il lavoro è completato quando sono definiti un ambito di streaming concordato e una configurazione che supporti il profiling headless di lunga durata.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- performance
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Tranquilla
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100