Add Perfetto trace output to profiling.sampling
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Python
- Estrellas
- 77.2k
- Forks
- 35.9k
- Métricas de merge de PR
- Métricas de PR pendientes
Descripción
Feature or enhancement
Proposal:
Feature
Add a Perfetto trace output backend to profiling.sampling (target: 3.16), alongside the existing other output formats. This would let users open Python sampling profiles directly in the Perfetto UI and analyse them with PerfettoSQL in the trace processor.
Opening this for design sanity-check before sending a PR, as suggested by @pablogsal in offline discussion. The most load-bearing question (how to emit protobuf without a runtime dependency) is covered below.
Motivation
- Perfetto is a trace viewer oriented systems-perf work, and being able to view Python samples there means a Python profile can sit on the same timeline as scheduler events, native CPU samples, and application instrumentation rather than living in a separate tool.
- Once trace processor can read the output, users get SQL-driven analysis of Python samples (top-N functions, per-thread breakdowns, custom aggregations).
Proposed design
Output format
Emit a Perfetto trace (perfetto.protos.Trace) containing:
- A
ProcessDescriptor/ThreadDescriptorper observed process/thread. InternedDatacarryingFrame,Mapping, andCallstackentries- One
TracePacketper sample containing aStackSamplemessage that references the interned callstack.
StackSample is part of a new set of public profiling protos I'm landing in Perfetto specifically so that producers like this one have a stable, transport-neutral surface to target (rather than reusing PerfSample, which is shaped by perf_event_open and leaks producer diagnostics into the data). The RFC is at https://github.com/google/perfetto/discussions/6027.
Concretely, each Python sample maps to roughly:
TracePacket {
timestamp: <ns>
trusted_packet_sequence_id: <seq>
interned_data { ... } // first packet only, or as new frames appear
stack_sample {
task_context_iid: <thread>
execution_context_iid: <cpu/mode> // optional
callstack_iid: <interned callstack>
primary_descriptor_iid: <"wall_time_ns" counter>
primary_weight: <ns since last sample>
}
}
Proto serialisation without a runtime dependency
The stdlib can't depend on protobuf, so I'd hand-roll the wire format for the specific message types we emit. This is tractable because:
- Proto wire format is tiny. It's varints + length-delimited + fixed32/64 + a tag byte per field. The whole encoder for the messages we need is on the order of a couple hundred lines of Python.
- We only encode, never decode. Decoding protobuf is signifcantly harder than encode.
- The set of message types is small and stable. I've specifically designated in the RFC upstream that these protos I'm adding are going to be "eternally stabe" protos which we won't change the wire format or semantics in a non-backwards compatible way.
- Precedent. Perfetto itself ships an ad-hoc proto encoder/decoder (
protozero) in C++ for similar reasons.
Sketch of the shape:
# Hand-rolled, no runtime deps.
def _varint(buf, n): ...
def _tag(buf, field_no, wire_type): ...
def _string(buf, field_no, s): ...
def _message(buf, field_no, payload): ...
def encode_stack_sample(buf, sample):
_tag(buf, 1, WIRE_VARINT); _varint(buf, sample.task_context_iid)
_tag(buf, 6, WIRE_VARINT); _varint(buf, sample.callstack_iid)
# ...
Field numbers and wire types come straight from the .proto definitions in the Perfetto RFC. Any wire type constants would be inlined as Python integers.
CLI surface
A new --format perfetto CLI flag to the profiling.sampling module in all the same places --gecko is allowed today.
Target
Python 3.16.
Prerequisites
- Perfetto RFC-0027: public stack-sampling and heap-profiling protos. Required before this lands so we're targeting the stable protos.
cc @pablogsal
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs
- gh-154541
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Comienza en la CLI profiling.sampling y sigue la ruta de salida existente de --gecko; después, lee el RFC de Perfetto enlazado sobre las definiciones de StackSample y de los proto relacionados. El trabajo estará terminado cuando --format perfetto emita un trace que se abra en la Perfetto UI y admita los datos de samples y threads de Python descritos.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- cli, performance
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 25/100