Repository navigation
An interpreter's initial thread can be accessed while finalizing #126914
Copy link
Copy link
Closed
Labels
3.12only security fixesonly security fixes3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)topic-subinterpreterstype-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump
Description
Activity
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)type-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump3.12only security fixesonly security fixes3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixes
on Nov 16, 2024 I think the easiest (and has the least amount of risk for backporting) is to just hold the runtime lock for all of thread state deletion.
Apparently, this is problematic for the free-threaded build, because QSBR and thread-local refcounting initialization cause a lock-ordering deadlock in some cases. So, instead, my solution is to just add a new field to the interpreter state that indicates when it's safe to access
_initial_thread, and that field is only used atomically (unlike something liketstate->_status). I'm hoping that should fix it without issues.- added a commit that references this issue
on Dec 2, 2024
Metadata
Metadata
Assignees
Labels
3.12only security fixesonly security fixes3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)topic-subinterpreterstype-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump
Projects
- StatusShow more project fieldsDone
Crash report
What happened?
While working on gh-126644, I came across an elusive bug that seems to occur when a lot of threads are trying to create a thread state for the same interpreter. Initially, I thought it was an issue with free-threading, but it occurs on the default build as well.
Here's a small reproducer:
As it turns out, this is because of this check: https://github.com/python/cpython/blob/main/Python/pystate.c#L1531
Inside
tstate_delete_common, the runtime lock is released after thethreads.headhas already been set toNULL, so other threads think erroneously think it's OK to try and use_initial_threadwhile it's still getting finalized. Possibly related,free_threadstatetries to reset_initial_threadback to the default settings without the runtime lock, which is possibly racy.There's a few ways to fix this, but I think the easiest (and has the least amount of risk for backporting) is to just hold the runtime lock for all of thread state deletion. This will introduce a very slight slowdown due to lock contention, but that should get better once #114940 lands. I have a PR ready.
CPython versions tested on:
3.12, 3.13, 3.14, CPython main branch
Operating systems tested on:
Linux
Output from running 'python -VV' on the command line:
No response
Linked PRs