Skip to content

pthread_exit & PyThread_exit_thread from PyEval_RestoreThread etc. are harmful #87135

Description

@gpshead
BPO 42969
Nosy @gpshead, @pitrou, @vstinner, @colesbury, @izbyshev, @jbms
PRs
  • gh-87135: Hang non-main threads that attempt to acquire the GIL during finalization #28525
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2021-01-19.19:35:46.531>
    labels = ['interpreter-core', 'type-bug', '3.9', '3.10', '3.11']
    title = 'pthread_exit & PyThread_exit_thread from PyEval_RestoreThread etc. are harmful'
    updated_at = <Date 2021-11-15.22:28:59.110>
    user = 'https://github.com/gpshead'

    bugs.python.org fields:

    activity = <Date 2021-11-15.22:28:59.110>
    actor = 'colesbury'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['Interpreter Core']
    creation = <Date 2021-01-19.19:35:46.531>
    creator = 'gregory.p.smith'
    dependencies = []
    files = []
    hgrepos = []
    issue_num = 42969
    keywords = ['patch', '3.2regression']
    message_count = 26.0
    messages = ['385288', '385289', '396669', '396679', '401886', '401922', '401958', '401959', '402218', '402219', '402220', '402264', '402266', '402474', '402475', '402487', '402509', '402519', '402556', '402558', '402559', '402560', '402571', '402608', '402684', '406363']
    nosy_count = 6.0
    nosy_names = ['gregory.p.smith', 'pitrou', 'vstinner', 'colesbury', 'izbyshev', 'jbms']
    pr_nums = ['28525']
    priority = 'normal'
    resolution = None
    stage = 'patch review'
    status = 'open'
    superseder = None
    type = 'behavior'
    url = 'https://bugs.python.org/issue42969'
    versions = ['Python 3.9', 'Python 3.10', 'Python 3.11']

    Linked PRs

    Activity

    1. gpshead commented on Jan 19, 2021

      @gpshead
      MemberAuthor

      BACKGROUND

      PyThread_exit_thread() calls pthread_exit() and is in turn called from a variety of APIs as documented in the C-API doc update from gh-80608.

      The pthread_exit() call was originally introduced as a way "resolve" a crashes from daemon threads during shutdown gh-46164. It did that. That fix to that even landed in 2.7.8 but was rolled back before 2.7.9 due to a bug in an existing application it exposed at the time (did we miss the implications of that? maybe). It remained in 3.x.

      PROBLEM

      pthread_exit() cannot be used blindly by any application. All code in the threaded application needs to be on board with it and always prepared for any API call they make to potentially lead to thread termination. Quoting a colleague: "pthread_exit() does not result in stack unwind or local variable destruction". This means that any code up the stack from the ultimate pthread_exit() call that has process-wide state implications that did not go out of its way to register cleanup with pthread_cleanup_push() could lead to deadlocks or lost resources. Something implausible to assume that code does.

      We're seeing this happen with C/C++ code. Our C++ builds with -fno-exceptions so uncatchable exception based stack unwinding as some pthread_exit implementations may trigger does not happen (and cannot be guaranteed anyways, see gh-87054). Said C/C++ code is calling back into Python from a thread and thus must use PyEval_RestoreThread() or similar APIs before performing Python C API calls. If the interpreter is being finalized from another thread... these enter a codepath that ultimately calls pthread_exit() leaving corrupt state in the process. In this case that unexpected thread simply disappearing can lead to a deadlock in our process.

      Fundamentally I do not believe the CPython VM should ever call pthread_exit() when non-CPython frames are anywhere in the C stack. This may mean we should never call pthread_exit() at all (unsure; but it'd be ideal).

      The documentation suggests that all callers in user code of the four C-APIs with the documented pthread_exit() caveats need auditing and pre-call _Py_IsFinalizing() API checks. But... I do not believe that would fix anything even if it were done. _Py_IsFinalizing() called without the GIL held means that it could change by the time the PyEval_RestoreThreads() API calls it internally do determine if it should exit the thread. Thus the race condition window would merely be narrowed, not eliminated. Not good enough.

      CURRENT WORKAROUND (Big Hammer)

      Change CPython to call abort() instead of pthread_exit() as that situation is unresolvable and the process dying is better than hanging, partially alive. That solution isn't friendly, but is better than being silent and allowing deadlock. A failing process is always better than a hung process, especially a partially hung process.

      SEMI RELATED WORK

      gh-87054 - appears to be avoiding some PyThread_exit_thread() calls to stop some crashes due to libgcc_s being loaded on demand upon thread exit.

    2. added
      interpreter-core(Objects, Python, Grammar, and Parser dirs)
      type-bugAn unexpected behavior, bug, or error
      on Jan 19, 2021
    3. gpshead commented on Jan 19, 2021

      @gpshead
      MemberAuthor

      C-APIs such as PyEval_RestoreThreads() are insufficient for the task they are asked to do. They return void, yet have a failure mode.

      They call pthread_exit() on failure today.

      Instead, they need to return an error to the calling application to indicate that "The Python runtime is no longer available."

      Callers need to act on that in whatever way is most appropriate to them.

    4. changed the title [-]pthread_exit & PyThread_exit_thread are harmful[/-] [+]pthread_exit & PyThread_exit_thread from PyEval_RestoreThread etc. are harmful[/+] on Jan 19, 2021
    5. vstinner commented on Jun 28, 2021

      @vstinner
      Member

      See also bpo-44434: "_thread module: Remove redundant PyThread_exit_thread() call to avoid glibc fatal error: libgcc_s.so.1 must be installed for pthread_cancel to work".

      New changeset 45a78f9 by Victor Stinner in branch 'main':
      bpo-44434: Don't call PyThread_exit_thread() explicitly (GH-26758)
      45a78f9

    6. vstinner commented on Jun 28, 2021

      @vstinner
      Member

      See also a discussion about the usefulness of daemon threads:
      python-trio/trio#2046

      I'm more in favor of deprecating daemon threads (in any interpreter, not only in subinterpreters). The current implementation is too fragile. There are still corner cases like the one described in this issue.

    7. jbms commented on Sep 15, 2021

      jbmsmannequin
      Mannequin

      Another possible resolution would to simply make threads that attempt to acquire the GIL after Python starts to finalize hang (i.e. sleep until the process exits). Since the GIL can never be acquired again, this is in some sense the simplest way to fulfill the contract. This also ensures that any data stored on the thread call stack and referenced from another thread remains valid. As long as nothing on the main thread blocks waiting for one of these hung threads, there won't be deadlock.

      I have a case right now where a background thread (created from C++, which is similar to a daemon Python thread) acquires the GIL, and calls "call_soon_threadsafe" on an asycnio event loop. I think that causes some Python code internally to release the GIL at some point, after triggering some code to run on the main thread which happens to cause the program to exit. While Py_FinalizeEx is running, the call to "call_soon_threadsafe" completes on the background thread, attempts to re-acquire the GIL, which triggers a call to pthread_exit. That unwinds the C++ stack, which results in a call to Py_DECREF without the GIL held, leading to a crash.

    8. vstinner commented on Sep 16, 2021

      @vstinner
      Member

      Change CPython to call abort() instead of pthread_exit() as that situation is unresolvable and the process dying is better than hanging, partially alive. That solution isn't friendly, but is better than being silent and allowing deadlock. A failing process is always better than a hung process, especially a partially hung process.

      The last time someone proposed to always call abort(), I proposed to add a hook instead: I added sys.unraisablehook. See bpo-36829.

      If we adopt this option, it can be a callback in C, something like: Py_SetThreadExitCallback(func) which would call func() rather than pthread_exit() in ceval.c.

      --

      Another option would be to add an option to disable daemon thread.

      concurrent.futures has been modified to no longer use daemon threads: bpo-39812.

      It is really hard to write a reliable implementation of daemon threads with Python subintepreters. See bpo-40234 "[subinterpreters] Disallow daemon threads in subinterpreters optionally".

      There is already a private flag for that in subinterpreters to disallow spawning processes or threads: an "isolated" subintepreter. Example with _thread.start_new_thread():

          PyInterpreterState *interp = _PyInterpreterState_GET();
          if (interp->config._isolated_interpreter) {
              PyErr_SetString(PyExc_RuntimeError,
                              "thread is not supported for isolated subinterpreters");
              return NULL;
          }

      Or os.fork():

          if (interp->config._isolated_interpreter) {
              PyErr_SetString(PyExc_RuntimeError,
                              "fork not supported for isolated subinterpreters");
              return NULL;
          }

      See also my article on fixing crashes with daemon threads:

    9. jbms commented on Sep 16, 2021

      jbmsmannequin
      Mannequin

      Regarding your suggestion of adding a hook like Py_SetThreadExitCallback, it seems like there are 4 plausible behaviors that such a callback may implement:

      1. Abort the process immediately with an error.

      2. Exit immediately with the original exit code specified by the user.

      3. Hang the thread.

      4. Attempt to unwind the thread, like pthread_exit, calling pthread thread cleanup functions and C++ destructors.

      5. Terminate the thread immediately without any cleanup or C++ destructor calls.

      The current behavior is (4) on POSIX platforms (pthread_exit), and (5) on Windows (_endthreadex).

      In general, achieving a clean shutdown will require the cooperation of all relevant code in the program, particularly code using the Python C API. Commonly the Python C API is used more by library code rather than application code, while it would presumably be the application that is responsible for setting this callback. Writing a library that supports multiple different thread shutdown behaviors would be particularly challenging.

      I think the callback is useful, but we would still need to discuss what the default behavior should be (hopefully different from the current behavior), and what guidance would be provided as far as what the callback is allowed to do.

      Option (1) is highly likely to result in a user-visible error --- a lot of Python programs that previously exited successfully will now, possibly only some of the time, exit with an error. The advantage is the user is alerted to the fact that some threads were not cleanly exited, but a lot of previously working code is now broken. This seems like a reasonable policy for a given application to impose (effectively requiring the use of an atexit handler to terminate all daemon threads), but does not seem like a reasonable default given the existing use of daemon threads.

      Option (2) would likely do the right thing in many cases, but main thread cleanup that was previously run would now be silently skipped. This again seems like a reasonable policy for a given application to impose, but does not seem like a reasonable default.

      Option (3) avoids the possibility of crashes and memory corruption. Since the thread stack remains allocated, any pointers to the thread stack held in global data structures or by other threads remain valid. There is a risk that the thread may be holding a lock, or otherwise block progress of the main thread, resulting in silent deadlock. That can be mitigated by registering an atexit handler.

      Option (4) in theory would allow cleanup handlers to be registered in order to avoid deadlock due to locks held. In practice, though, it causes a lot of problems:

      • The CPython codebase itself contains no such cleanup handlers, and I expect the vast majority of existing C extensionss are also not designed to properly handle the stack unwind triggered by pthread_exit. Without proper cleanup handlers, this option reverts to option (5), where there is a risk of memory corruption due to other threads accessing pointers to the freed thread stack. There is also the same risk of deadlock as in option (3).
      • Stack unwinding interacts particularly badly with common C++ usage because the very first thing most people want to do when using the Python C API from C++ is create a "smart pointer" type for holding a PyObject pointer that handles the reference counting automatically (calls Py_INCREF when copied, Py_DECREF in the destructor). When the stack unwinds due to pthread_exit, the current thread will NOT hold the GIL, and these Py_DECREF calls result in a crash / memory corruption. We would need to either create a new finalizing-safe version of Py_DECREF, that is a noop when called from a non-main thread if _Py_IsFinalizing() is true (and then existing C++ libraries like pybind11 would need to be changed to use it), or modify the existing Py_DECREF to always have that additional check. Other calls to Python C APIs in destructors are also common.
      • When writing code that attempts to be safe in the presence of stack unwinding due to pthread_exit, it is not merely explicitly GIL-related calls that are a concern. Virtually any Python C API function can transitively release and acquire the GIL and therefore you must defend against unwind from virtually all Python C API functions.
      • Some C++ functions in the call stack may unintentionally catch the exception thrown by pthread_exit and then return normally. If they return back to a CPython stack frame, memory corruption/crashing is likely.
      • Alternatively, some C++ functions in the call stack may be marked noexcept. If the unwinding reaches such a function, then we end up with option (1).
      • In general this option seems to require auditing and fixing a very large amount of existing code, and introduces a lot of complexity. For that reasons, I think this option should be avoided. Even on a per-application basis, this option should not be used because it requires that every C extension specifically support it.

      Option (5) has the risk of memory corruption due to other threads accessing pointers to the freed thread stack. There is also the same risk of deadlock as in option (3). It avoids the problem of calls to Python C APIs in C++ destructors. I would consider this options strictly worse than option (3), since there is the same risk of deadlock, but the additional risk of memory corruption. We free the thread stack slightly sooner, but since the program is exiting soon anyway that is not really advantageous.

      The fact that the current behavior differs between POSIX and Windows is particularly unfortunate.

      I would strongly urge that the default behavior be changed to (3). If Py_SetThreadExitCallback is added, the documentation could indicate that the callback is allowed to terminate the process or hang, but must not attempt to terminate the thread.

    10. jbms commented on Sep 16, 2021

      jbmsmannequin
      Mannequin

      Regarding your suggestion of banning daemon threads: I happened to come across this bug not because of daemon threads but because of threads started by C++ code directly that call into Python APIs. The solution I am planning to implement is to add an atexit handler to prevent this problem.

      I do think it is reasonable to suggest that users should ensure daemon threads are exited cleanly via an atexit handler. However, in some cases that may be challenging to implement, and there is also the issue of backward compatibility.

    11. vstinner commented on Sep 20, 2021

      @vstinner
      Member

      PyThread_exit_thread() is exposed as _thread.exit() and _thread.exit_thread().

      PyThread_exit_thread() is only called in take_gil() (at 3 places in the function) if tstate_must_exit(tstate) is true. It happens in two cases:

      • (by design) at Python exit if a daemon thread tries to "take the GIL": PyThread_exit_thread() is called.

      • (under an user action) at Python exit if threading._shutdown() is interrupted by CTRL+C: Python (regular) threads will continue to run while Py_Finalize() is running. In this case, when a (regular) thread tries to "take the GIL", PyThread_exit_thread() is called.

    12. vstinner commented on Sep 20, 2021

      @vstinner
      Member

      I don't think that there is a "good default behavior" where Python currently calls PyThread_exit_thread().

      IMO we should take the problem from the other side and tries to reduce cases when Python can reach this case. Or even make it impossible if possible. For example, *removing* daemon threads would remove the most common case when Python has to call PyThread_exit_thread().

      I'm not sure how to make this case less likely when threading._shutdown() is interrupted by CTRL+C. This function can likely hang if a thread is stuck for whatever reason. It's important than an user is able to "interrupt" or kill a stuck process with CTRL+C (SIGINT). It's a common expectation on Unix, at least for me.

      Maybe threading._shutdown() should be less nice and call os._exit() in this case: exit *immediately* the process in this case. Or Python should restore the default SIGINT handler: on Unix, the default SIGINT handler immediately terminate the process (like os._exit() does).

      I don't think that abort() should be called here (raise SIGABRT signal), since the intent of an user pressing CTRL+C is to silently terminate the process. It's not an application bug, but an user action.

    13. vstinner commented on Sep 20, 2021

      @vstinner
      Member

      See also bpo-13077 "Windows: Unclear behavior of daemon threads on main thread exit".

    14. 55 remaining items

    15. encukou commented on Jun 26, 2025

      @encukou
      Member

      PR for PythonFinalizationError on unacquirable threading.Lock: #135991

    16. added a commit that references this issue on Jul 1, 2025
    17. encukou commented on Jul 1, 2025

      @encukou
      Member

      I think this issue can be closed now.
      Daemon threads hang on finalization, joining them gives PythonFinalizationError; that's about as much as we can do with current API.
      PEP-788 proposes better API for native callbacks and destructors.

    18. added a commit that references this issue on Jul 11, 2025
    19. added a commit that references this issue on Jul 12, 2025
    20. added a commit that references this issue on Jul 13, 2025
    21. added a commit that references this issue on Aug 4, 2025
    22. added 2 commits that reference this issue on Aug 15, 2025
    23. added a commit that references this issue on Aug 19, 2025
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    Labels

    3.11only security fixes3.12only security fixes3.13only security fixes3.14bugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)type-bugAn unexpected behavior, bug, or error

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions