Repository navigation
Use dual stacks to separate control and data #157559
Description
Activity
@pablogsal would this work for Tachyon?
This would also make implementing https://discuss.python.org/t/show-builtin-functions-and-classes-in-tracebacks-and-profiles/24866 much simpler.
- addedtype-featureA feature request or enhancementA feature request or enhancementinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Sep 15, 2026 Tachyon is advertised as zero-overhead, but it isn't. It requires extra code on the fast path of yields and returns and couples the VM and profiler in undocumented
I’m happy to discuss the stack split, but I think this paragraph is a gross mischaracterization of Tachyon and the constraints involved.
The bookkeeping on yields and returns has been measured to be below noise several times and the fact that you cannot do arbitrary changes to the VM doesn't allow you to say the profiler is not zero overhead.and hard to maintain
I respectfully disagree with this but of course you can disagree with me :)
Any changes to the base pointer changes will be protected by a memory fence, so other processors will see the change.
A fence alone doesn’t specify how an external reader obtains consistent base and top pointers, or how long the allocation remains readable. What happens if the stack moves between reading those pointers and copying the frames? Can the reader observe pointers from different allocations? We need allocation lifetime and publication rules, plus a way to detect an inconsistent read and retry.
The control frame pointer will not be synchronized, so may appear out-of-date to other processors.
What does a stale pointer allow us to observe? Frames might have returned, their slots might have been reused, and their executable objects might no longer be alive. Approximate samples are expected, but we need to distinguish an outdated instruction position from a record containing fields from different invocations. The contract should explain which inconsistencies are possible and what readers can validate.
PyObject *executable; // The code object, or non-Python callable, for this frame.
This is useful, but profilers combining Python and native stacks also need to associate interpreter entries with positions on the native stack. Currently, interpreter entry-frame addresses can provide that association. If those records move into a separate array, we need an equivalent native-stack anchor or mapping. Recording a non-Python callable doesn’t by itself tell us where arbitrary native frames belong relative to Python frames.
_PyInterpreterDataFrame *framepointer;
This looks like a suitable way to preserve access to locals and arguments, but it appears only in the implementation view, not the new profiler contract. Profilers that inspect values need a discoverable mapping from a control record to its data frame and the relevant data offsets.
These frames are mostly allocated in large chunks, but may not be contiguous as generator frames are allocated as part of the generator.
How will suspended generators and coroutines be represented after the split? Their saved execution state isn’t part of an active thread’s control stack, but external tools still need to inspect it and follow await relationships. Tachyon already uses that information for async inspection. We need to specify where the saved control state lives and how readers reach it, including during suspension and resumption.
_Py_CODEUNIT instr_ptr; / Instruction currently executing (may be approximate) */
For free-threaded builds, interpreting this pointer requires accounting for thread-local bytecode. The internal structure mentions tlbc_index, but the profiler view doesn’t expose it. We need that information or an equivalent normalized bytecode offset to preserve line and opcode attribution. It would also help to define what “approximate” means, including when the stored position is updated.
It requires extra code on the fast path of yields and returns
I’m not willing to remove the frame-cache bookkeeping as part of this change. It provides the validity information profilers need to reuse previously sampled frames, and its runtime cost has been measured to be below noise several times. Fixed-size control frames don’t replace that information: slots can still be reused by different invocations. We can adapt the bookkeeping to the new layout, but removing it without an equivalent mechanism would regress existing profiling functionality and performance.
To be absolutely clear on what are the constraints here. The new layout needs to preserve:
- Rules for reading the stack while it changes, including detecting invalid reads and keeping allocations readable.
- The frame-cache bookkeeping, or an equivalent way to safely reuse cached frames.
- A way to match Python frames to their positions in the native stack. Currently profilers use the entry frame information and the addresses of the stack allocated entry frame.
- Access to locals and arguments.
- Access to suspended generator and coroutine state, including await relationships.
- The information needed to resolve line numbers and bytecode offsets in free-threaded builds.
We can change the layout, but these features need to keep working.
A way to match Python frames to their positions in the native stack. Currently profilers use the entry frame information and the addresses of the stack allocated entry frame.
A way to achueve this with the proposed layout would be to keep an interpreter-entry marker in the control stack, with an explicit address pointing into the corresponding native stack frame. That preserves what some profilers needs without depending on where control records are allocated. Today, we creates a local
_PyEntryFrameentry in the evaluator and profilers uses its address to find the native frame whose stack-pointer range contains it. It doesn’t need the whole entry structure for that mapping, just an address with the same lifetime and location. With the new layout perhaps we could do the following:- On entry from C into the evaluator, push a control record marked as an interpreter entry.
- Store a
native_stack_anchorin that record. It points to a small local marker in the evaluator’s native stack frame, or a suitable existing local. - Keep that marker alive for the entire evaluator invocation, including execution through tail-call interpreter helpers.
- Remove the control record before the evaluator invocation returns.
- Expose the entry-record kind and anchor offset through debug offsets.
Normal Python-to-Python calls wouldn’t need their own anchor. Nested calls through C would create another entry marker:
Control stack Native stack Python B evaluator invocation 2 Entry marker → anchor 2 C extension callback Python A evaluator invocation 1 Entry marker → anchor 1 C callerThen tools would read the anchor instead of using the control record’s own address. Its existing stack-pointer-range matching could then stay largely unchanged.
Ok, I have dedicated some time to think about this. If you want to keep thinking about it I would suggest the following adaptations. These would let the VM change the object-stack layout while keeping the information that profilers need in the control stack. I would use your control-frame structure, with one change:
typedef struct { PyObject *executable; _Py_CODEUNIT *instr_ptr; union { _PyInterpreterDataFrame *framepointer; uintptr_t native_stack_anchor; /* Interpreter entry only. */ }; _PyStackRef *stackpointer; uint8_t owner; uint8_t visited; uint16_t return_offset; #ifdef Py_GIL_DISABLED int32_t tlbc_index; #endif } _PyControlFrame;
For interpreter-entry records, the union holds an address inside the corresponding native activation. The address stays valid until that evaluator invocation returns. Profilers use it to place Python frames among native frames; they do not dereference it. This replaces the association currently provided by the entry frame’s own address. It adds no space to your proposed structure and needs no separate list or additional remote read.
I would expose the offsets of these fields through the debug metadata, including
ownerandtlbc_index. Profilers could then copy the control array and obtain the information needed for Python stacks, native-stack merging, and free-threaded instruction positions from that copy. Basic sampling would not need to read the object stack.I would also keep the frame-cache bookkeeping. We could change the bookmark from a frame address to a control-stack index so that it remains valid across relocation:
typedef struct { size_t frame_index; /* SIZE_MAX means no bookmark. */ uint64_t pop_sequence; } _PyProfilerBookmark;
This would retain the current ability to validate and reuse cached callers instead of copying and processing the full stack on every sample. The update rules would still cover returns, yields, and exception unwinding.
We still have problems with stack relocation and lifetime and the fence doesn't help for external profilers. One option here is to retain old allocations until thread-state destruction and use a relocation sequence to detect changes during a read. Keep that sequence, the base, the top index, and the bookmark together so readers can check them in the same metadata reads. Concurrent frame updates still need defined validation rules as a relocation sequence alone does not make the stack copy consistent.
For suspended generators and coroutines, we need to retain an embedded saved control record and expose its offset. This gives external tools direct access to saved execution state without searching active thread stacks or following another allocation. Access to await relationships also needs to remain available.
The data-frame layout can stay as you propose. Tools that need locals or arguments can use
framepointer, while ordinary sampling stays independent of that layout. We should check the read cost for value inspection, since some values currently arrive with the copied frame buffer.This keeps the fixed-size control stack and gives the VM freedom to arrange the object stack. It also preserves cached sampling and keeps native-stack information in the same copy. I would verify remote-read counts, copied bytes, and sampling time for these paths before landing, so we can catch any cost introduced by the split.
Thanks for the feedback @pablogsal
I'd like to distill this down to some set of fixed requirements that we can design an interface around.
So, I have a few questions:A fence alone doesn’t specify how an external reader obtains consistent base and top pointers, or how long the allocation remains readable. What happens if the stack moves between reading those pointers and copying the frames? Can the reader observe pointers from different allocations? We need allocation lifetime and publication rules, plus a way to detect an inconsistent read and retry.
Aren't those already issues with the current implementation?
_PyThreadState_UpdateLastProfiledFrame()doesn't include any synchronization, so other threads/processes might see stale values.What does a stale pointer allow us to observe? Frames might have returned, their slots might have been reused, and their executable objects might no longer be alive
Again, isn't this true for the current implementation?
A way to achieve this with the proposed layout would be to keep an interpreter-entry marker in the control stack, with an explicit address pointing into the corresponding native stack frame.
Would a
NULLpointer be sufficient as a marker for the entry frame, or does it need to point into the C stack?
If it does need to point into the C stack, does it need to point at anything in particular?I’m not willing to remove the frame-cache bookkeeping as part of this change. It provides the validity information profilers need to reuse previously sampled frames, and its runtime cost has been measured to be below noise several times. Fixed-size control frames don’t replace that information: slots can still be reused by different invocations. We can adapt the bookkeeping to the new layout, but removing it without an equivalent mechanism would regress existing profiling functionality and performance.
I’m not willing to remove the frame-cache bookkeeping as part of this change
Does that work now?
What guarantee is there that the bookkeeping information is any more up to date than the stack? Both are unsynchronized and open to a variety of race conditions.We still have problems with stack relocation and and the fence doesn't help for external profilers.
Relocations shouldn't be a problem. The memory fence means that you can double-check correctly across processes.
An out-of-date view of the stack is still going to be problem, but you must have mechanisms for dealing with that now. Why would those mechanisms not continue to work?
For suspended generators and coroutines, we need to retain an embedded saved control record and expose its offset. This gives external tools direct access to saved execution state without searching active thread stacks or following another allocation. Access to await relationships also needs to remain available.
Can you explain what information you need here?
Do you want to recreate the implicit task that a chain of suspended coroutines represents?Suspended generators would have no control record, but the information is already present in the generator, and we won't be removing it.
If I'm understanding what you need here, I think the best approach here might be to be more explicit about the concept of a subgenerator/awaited coroutine and when aYIELD_VALUEhappens as part of ayield fromorawaitmove the subgenerator off the evaluation stack and into a field of the caller generator/coroutine, although this doesn't seem directly relevant to this proposal.One thing I'd like to point out:
The less information in the control frame, the smaller they become and the cheaper they are to copy. If we can keep them down to 32 bytes, a single 4kb page will contain 128 frames which is enough for the whole stack in many cases, in which case copying the whole stack might be cheaper than having to do many cross-process reads and complex checking of state.
The internal structure mentions tlbc_index, but the profiler view doesn’t expose it. We need that information or an equivalent normalized bytecode offset
The
tlbcstuff is a bit hacky, so is likely to change. I'd rather just dedicate an opaque field for information to derive the bytecode offset from and a provide a helper function, that we can have tests for, to parse that information.Thanks for your questions. Let me try to answer them as clearly as I can. We can also discuss about the details together at the sprint.
Aren't those already issues with the current implementation?
Yes, but those races are expected and acceptable for sampling profilers. By default, Tachyon reads memory while the application keeps running. Unless the user selects
--blocking, it does not stop the application to obtain a consistant snapshot.The requirement is to detect inconsistent data and respond. The reader can reject a cache entry, continue walking frames, or drop the sample. We do not need to eliminate every race. We do need to preserve this information that makes these checks possible.
The current reader checks frame ownership, limits stack walks, rejects incomplete walks, and validates cache entries. These checks do not detect every possible inconsistency. The same limitation applies to stale frame pointers and executable lifetimes.
I am not asking the new layout to solve every existing race. I am asking which checks remain valid and how readers should handle changes to the new stack.
Would a NULL pointer be sufficient as a marker for the entry frame, or does it need to point into the C stack?
It needs to point into the C stack frame for that evaluator call.
NULLidentifies an entry boundary, but does not show where it belongs in the native stack.For example:
Python A → C extension → Python B → another native functionProfilers and debuggers that require merging the stacks need to recover this order. They compare the entry address with the address ranges of native stack frames. This also allows to distinguish nested calls to the evaluator.
The address does not need to point to particular data. The reader does not dereference it. A local variable can provide the address, but it must remain valid for the whole evaluator call. This must also work with the tail-call interpreter.
We can change how we store that address. We cannot replace it with
NULLwithout another way to locate the Python frames in the native stack.Does that work now?
Yes. It provides cache validation, not an atomic snapshot. There is no separate guarantee that the bookkeeping values are newer than the stack contents.
Its purpose is to record changes that the current stack cannot show:
def caller(): worker() # First sample occurs inside this call. worker() # Second sample occurs inside this call.
Both calls to
worker()can use the same frame address or array slot. They can also have the same stack depth and code objects. But the caller’s position has changed. Reusing the old caller position would attribute the second sample to the first call.When the bookmarked frame leaves the stack, the VM moves the bookmark toward its caller and increments the sequence. Readers can detect that change when they observe the update. Comparing addresses alone cannot give us this informations.
The cache reader checks the saved address and sequence, refreshes the current frame, and reads the bookmark again. Only then does it accept cached callers. Partial hits also check whether the sequence change matches the expected number of removed frames.
These checks do not eliminate concurrent-read races. They still detect changes that readers would otherwise miss. They also matter with
--blocking, because frames can return and be replaced between samples.On your suggestion to copy a small control stack instead of using the cache: that is worth meassuring. Smaller records should improve full walks. But the cache avoids both reading unchanged callers and processing their records again:
main → A → B → C Sample 1: read and decode the stack. Sample 2: refresh C, validate the bookmark, reuse main/A/B. Sample 3: refresh C, validate the bookmark, reuse main/A/B.The current full-hit path does not copy stack chunks. It reuses decoded callers. It still assembles the output, so I am not claiming the entire sample has a fixed cost.
A page copy might beat some combinations of separate reads and checks. But copying is only a part of the cost. Readers must also process the records, and deep stacks can exceed one page.
This is a crucial optimization for Tachyon and other profilers, especially at high sampling rates and across many threads. I am not willing to remove the bookkeeping based on an assumed improvement. Its VM cost has been measured below noise several times.
We can adapt the bookkeeping to the new layout. Compact records and caching can coexist. We should compare full walks, full hits, and partial hits before deciding otherwise.
Relocations shouldn't be a problem.
We can handle relocation. My statement that the fence “doesn’t help” was a bit too broad. It can help order publication, but we still need the reader protocol.
For example:
Reader: reads the old base address. Target: moves the stack and releases the old allocation. Reader: tries to copy from the old address.What does the reader check before and after copying? When can the old allocation be freed or reused? An unreadable address produces an error that we can handle. A reused allocation can still produce a successful read, so that case also needs consideration.
On Linux,
process_vm_readv()does not guarantee atomic transfers. A fence in the target does not, by itself, specify how the reader validates the copied data.Retaining allocations and adding a generation were possible solutions. I do not require those particular solutons. Existing error handling can remain, but bookmarks and validation checks need defined behaviour across relocation.
Can you explain what information you need here?
Yes, we need to reconstruct the suspended coroutine chain. We need each coroutine’s code, saved instruction position, execution state, and awaited or delegated object.
Tachyon reads this information in parse_coro_chain(). It finds the next awaited object through the saved evaluation stack in handle_yield_from_frame(). It separately follows task
awaited_byrelationships.We do not require an embedded control record if readers can still access this information in the generator. Your proposed field for the awaited object would simplify the reader. It would also make the reader depend less of the evaluation-stack layout.
I agree that field can be a separate change. This proposal needs to preserve access to the saved state and relationships.
On your suggestion to use an opaque field and a tested helper for bytecode positions: that can work. We need to resolve the position, not preserve
tlbc_indexitself.The helper must run in the reader process without executing code in the target. It should use copied control records and metadata that readers can cache. It must not require a new remote lookup or syscall for every frame just to decode that field.
If it needs remote metadata, the interface should support caching and batched reads. It also needs a way to detect stale metadata. We should check those costs against the current implementation. Hiding the format behind a helper must not make profiling slower through additional syscalls.
External profilers and debuggers also need enough versioned information to use or implement the decoder. Tests for that interface would be usefull.
Finally, tools that inspect locals and arguments need a way to find the data frame. Basic stack sampling should not require those reads but some advance samplers need it for example to resolve importlib frames and map it to the imported modules via the first argument to those known functions. We should measure value inspection separately, since the split can add reads for values that currently arrive with the frame buffer.
@markshannon apart from the profiling concerns, I think we should not have any non localsplus items in the data stack.
The main reasoning is to allow for zero-copy Python-to-Python calls in the specializer.
E.g. if you have the data stack overlap, then the current caller's stack automatically become the callee's locals. That will allow zero-copy Python calls even in the specializing interpreter without a JIT.
Currently, in CPython, the stack is implemented as a linked list of frames. These frames are mostly allocated in large chunks, but may not be contiguous as generator frames are allocated as part of the generator.
Each frame's size depends on its code object, which makes scanning the scan slow and requires dynamic memory allocation to account for varying stack sizes, making it hard for external profilers, like Tachyon, to snapshot the stack without relying on undocumented assumptions about the VM.
Tachyon is advertised as zero-overhead, but it isn't. It requires extra code on the fast path of yields and returns and couples the VM and profiler in undocumented and hard to maintain ways preventing improvements in the VM e.g. #148681
By splitting the stack into two parts, a control stack and an object stack, we can make the control stack simpler, smaller and more regular, which would:
Performance impact
Having two stacks, means two stack pointers, using a register in the interpreter and JIT.
However, having a control stack of fixed-sized control frames, would simplify recursion depth checking and other bookkeeping tasks.
I expect the additional costs and the savings to largely cancel out.
Primarily, this is about decoupling profilers from the VM design, not performance, so a small initial slowdown would be acceptable.
Having the object stack being purely composed of object pointers, and no control, may allow some additional optimizations across calls, slightly reduced stack memory use, and slightly faster stack scanning for the GC, so this might eventually give a small performance boost.
Control frames
Profiler view
To a profiler, or other out-of-process tool, a control frame will look like this:
If callable is a code object, then additional information is available:
Implementation
The control frame contains all the data that isn't object pointers from the current interpreter frame, plus the pointer to the code object:
Debug info and guarantees for profilers
Out of process debuggers and profilers require information to traverse internal data structures. The VM will provide this information:
uint32_t control_frame_offsetoffset of the control frame pointer in the thread stateuint32_t control_base_offsetoffset of the pointer to the base of the stack in the thread stateuint32_t control_frame_sizethe size (in bytes) of a control frame ==CONTROL_FRAME_SIZEabove.Any changes to the base pointer changes will be protected by a memory fence, so other processors will see the change.
The control frame pointer will not be synchronized, so may appear out-of-date to other processors.
Prior discussion focused on implementation: faster-cpython/ideas#675
#115946 explains why this helps profilers.