Repository navigation
Add Perfetto trace output to profiling.sampling #150327
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on May 23, 2026 To clarify, I'm interested in writing a PR for this but I wanted to confirm the design was good before I did so!
I'm also interested in adopting
PerfettoWriter. I'm currently working ongc-monitor, which outputs Chrome Trace JSON usingX,C,M, andIevent types. I'd like to switch toPerfettoWriterto stream profiling data directly, with periodic flushes to disk if possible.- addedstdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory
on May 24, 2026 More generally, can't we actually design something where users can register their own format instead of having it in
profiling.sampling? it would be easier to have PyPI plugins rather than having everything in the stdlib.Reacted by Maxim Martynov and Eduardo Villalpando Mellocc @ivonastojanovic @pablogsal for thoughts on a new design (or if something was already decided); we do have JSONL output now I think though, if it helps, so I think we already have most of the machinery in place but we could have a more "native" way to plug-in formatters (in a way, we're just providing a publish-subscribe pattern and consumers just need to attach observers which are responsible to format whatever they are given and export it to other consumers).
I agree, I think we should consider exposing a public API for this as that would make it much simpler to add new output types and those could also evolve faster than the ones in the stdlib. We specifically avoided doing this in the initial implementation to avoid having to worry about API design too much, but now that Tachyon is pretty feature complete, it would be a good time to provide such an API.
The
Collectorinterface is pretty simple, we just need to to offer a way to register new collectors. That could be a new CLI arg to specify apackage.module:functionas a factory function to instantiate the collector (e.g.--collector mypkg.mycollector:func. We should create a separate issue for this, but I'm curious what @pablogsal and @ivonastojanovic thinks.cc @ivonastojanovic @pablogsal for thoughts on a new design (or if something was already decided); we do have JSONL output now I think though, if it helps, so I think we already have most of the machinery in place but we could have a more "native" way to plug-in formatters (in a way, we're just providing a publish-subscribe pattern and consumers just need to attach observers which are responsible to format whatever they are given and export it to other consumers).
I agree, I think we should consider exposing a public API for this as that would make it much simpler to add new output types and those could also evolve faster than the ones in the stdlib. We specifically avoided doing this in the initial implementation to avoid having to worry about API design too much, but now that Tachyon is pretty feature complete, it would be a good time to provide such an API.
The
Collectorinterface is pretty simple, we just need to to offer a way to register new collectors. That could be a new CLI arg to specify apackage.module:functionas a factory function to instantiate the collector (e.g.--collector mypkg.mycollector:func. We should create a separate issue for this, but I'm curious what @pablogsal and @ivonastojanovic thinks.I like the idea, I have a few initial thoughts on the design, but we can dive into the details once we open an issue. A few thoughts on the design:
- I like the
--collector module:factoryapproach. Correct me if I'm wrong, but the factory just needs to return something that implementscollect()andexport()? We should think about which args the factory receives and which ones should be mandatory vs optional. - We should probably also export
StackTraceCollectoralongsideCollector. It already handles thread iteration, frame filtering, so custom collectors can focus on processing stacks instead of reimplementing the collection logic. - We might also want to clean up
_create_collector()and the output handling as part of this. There's currently a fair amount of format-specific branching (filename generation, extension mapping, etc.) that could disappear if each collector is responsible for its own construction via a factory.
- I like the
So is the consensus here that we should close this issue and instead design a generic export API? Personally I think it would be nice to not have to install another library to get the Perfetto output but totally understand if we want to go another way!
I would love if @pablogsal could chip in as well because we discussed a bit offline on this which led to me opening this issue in the first place.
Thanks @LalitMaganti, I think we should have perfetto by default and we can do the generic API later. I think the format is useful enough that unless is too complex or a lot of code we should have it built in. The only thing I'd want confirmed before the PR lands is that RFC-0027 is sufficiently stable upstream so we're not chasing a proto reshape in 3.16.
Thanks @LalitMaganti, I think we should have perfetto by default and we can do the generic API later. I think the format is useful enough that unless is too complex or a lot of code we should have it built in.
Thanks for chiminng in! I think it should be quite simple PR for what we need :)
The only thing I'd want confirmed before the PR lands is that RFC-0027 is sufficiently stable upstream so we're not chasing a proto reshape in 3.16.
So the top level protos have not gone in yet (the ones which provide the timestamp, indicate the leaf of the callstack etc) but for the actual callstack/callframe protos we'll be reusing the ones which Perfetto has had for ~7 years so they will certainly not be changing :)
I'll send a PR soonish, hopefully later this week!
It too me a bit of time to stabilize the upstream protos and get all the code in; I didn't want to send anything before I had that done! But now all of that is complete.
I've sent #154541 for fixing this issue, it's my first time contributing so please excuse any mistakes I made!
Also while I was building this, one thing I noticed is that the binary format does not appear to include the pid of the process being profiled somewhere easily accessible. I found I could reach into some of the interpreter structs to figure out but this felt like a hack. Was this an intentional choice or an oversight?
The reason I ask is that when someone does binary capture and later converts to another format (e.g. Gecko or Perfetto), right now the pid used is the pid of the profiler process not the process being profiled. While this sort of doesn't matter if you are looking at the profile alone, Perfetto supports merging multiple, indepdently collected profiles together on a timeline (e.g. Python CPU profiling + scheduler tracing). For that, we do need acutally accurate pids.
I can file another issue if we should discuss this elsewhere?
The reason I ask is that when someone does binary capture and later converts to another format (e.g. Gecko or Perfetto), right now the pid used is the pid of the profiler process not the process being profiled. While this sort of doesn't matter if you are looking at the profile alone, Perfetto supports merging multiple, indepdently collected profiles together on a timeline (e.g. Python CPU profiling + scheduler tracing). For that, we do need acutally accurate pids.
I can file another issue if we should discuss this elsewhere?
There's more places where collectors (especially honest diffs and replays) yearn for more data in the binary format: #154105 #154101. I posted a comment #154105 (review) and my belief is that we should consider some flexible metadata struct (ideally: ready for streaming but that's easy since the size is known in advance) in the binary format, to avoid constant Whac-A-Mole.
Specifically, there's also more places where we describe nothing or always the host, not the target:
"processName": "Python Process", cpython/Lib/profiling/sampling/gecko_collector.py
Lines 827 to 830 in 66e313f
"abi": platform.machine(), "misc": "Python profiler", "oscpu": platform.machine(), "platform": platform.system(), cpython/Lib/profiling/sampling/gecko_collector.py
Lines 838 to 839 in 66e313f
"physicalCPUs": os.cpu_count() or 0, "logicalCPUs": os.cpu_count() or 0, cpython/Lib/profiling/sampling/heatmap_collector.py
Lines 404 to 406 in 66e313f
"python_version": sys.version, "python_implementation": platform.python_implementation(), "platform": platform.platform(), "<!-- PYTHON_VERSION -->", f"{sys.version_info.major}.{sys.version_info.minor}"
Feature or enhancement
Proposal:
Feature
Add a Perfetto trace output backend to
profiling.sampling(target: 3.16), alongside the existing other output formats. This would let users open Python sampling profiles directly in the Perfetto UI and analyse them with PerfettoSQL in the trace processor.Opening this for design sanity-check before sending a PR, as suggested by @pablogsal in offline discussion. The most load-bearing question (how to emit protobuf without a runtime dependency) is covered below.
Motivation
Proposed design
Output format
Emit a Perfetto trace (
perfetto.protos.Trace) containing:ProcessDescriptor/ThreadDescriptorper observed process/thread.InternedDatacarryingFrame,Mapping, andCallstackentriesTracePacketper sample containing aStackSamplemessage that references the interned callstack.StackSampleis part of a new set of public profiling protos I'm landing in Perfetto specifically so that producers like this one have a stable, transport-neutral surface to target (rather than reusingPerfSample, which is shaped byperf_event_openand leaks producer diagnostics into the data). The RFC is at google/perfetto#6027.Concretely, each Python sample maps to roughly:
Proto serialisation without a runtime dependency
The stdlib can't depend on
protobuf, so I'd hand-roll the wire format for the specific message types we emit. This is tractable because:protozero) in C++ for similar reasons.Sketch of the shape:
Field numbers and wire types come straight from the
.protodefinitions in the Perfetto RFC. Any wire type constants would be inlined as Python integers.CLI surface
A new
--format perfettoCLI flag to the profiling.sampling module in all the same places--geckois allowed today.Target
Python 3.16.
Prerequisites
cc @pablogsal
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs