Skip to content

ClosedResourceError + 'Unexpected ASGI message ... after response already completed' race in _handle_post_request notification path (streamable-http) #3641

Description

@djbclark

Bug: RuntimeError: Unexpected ASGI message 'http.response.start' sent, after response already completed following ClosedResourceError in _handle_post_request notification/response path

Environment

  • mcp: 2.1.1
  • starlette: 1.6.0
  • uvicorn: 0.52.4
  • anyio: 4.14.2
  • Python: 3.13.14
  • OS: macOS 27.0.0 (Darwin, arm64)
  • Server: basic-memory 0.23.2 (FastMCP-based), streamable-http transport, stateful (not stateless_http), single uvicorn worker, run as a long-lived local service (launchd) on 127.0.0.1.

Description

Our server logs this pair of exceptions every 1-2 hours of normal multi-client usage (several unrelated local MCP clients making tool calls concurrently against the same long-running session). Each occurrence is almost always transient — the server keeps serving the next request fine within milliseconds — but on one occasion it coincided with the server task group getting wedged: the process stayed alive and listening, but stopped responding to any further request, accumulating CPU in a runnable state until we force-restarted it. I can't yet prove the wedge and this exception pair are causally linked (vs. coincidental timing), but I haven't found another candidate in the logs for that incident, so I'm reporting the reproducible half of it (the exception pair) and flagging the wedge as a possible consequence.

The two paired exceptions (always appear together, in this order)

1. In _handle_post_request's notification/response branch (around streamable_http.py:588, await writer.send(session_message)):

Error handling POST request
Traceback (most recent call last):
  File ".../mcp/server/streamable_http.py", line 588, in _handle_post_request
    await writer.send(session_message)
    ...
    raise ClosedResourceError
anyio.ClosedResourceError

2. Immediately after (same request, different connection, usually 1-4 seconds later):

ERROR:    Exception in ASGI application
Traceback (most recent call last):
  File ".../mcp/server/streamable_http.py", line 588, in _handle_post_request
    await writer.send(session_message)
    ...
    raise RuntimeError(f"Unexpected ASGI message '{message['type']}' sent, after response already completed.")
RuntimeError: Unexpected ASGI message 'http.response.start' sent, after response already completed.

Where this sits relative to known issues

This looks related to, but distinct from, the previously-reported/fixed races in this area:

Ours happens on the notification/response branch at line 588 — the code path that:

  1. Immediately sends a 202 Accepted HTTP response for a non-JSONRPCRequest message (notification or response), then
  2. Forwards the message to writer for the session's message router to process.

The second exception (Unexpected ASGI message 'http.response.start' sent, after response already completed) implies something tried to start an ASGI response a second time on a connection that had already fully completed its response — i.e., there appear to be two handlers (or one handler re-entered) racing on the same underlying ASGI send callable, not just a stream that was already closed.

Hypothesis

With several concurrent clients attached to the same session, two POST requests can race such that:

  • Request A completes its 202 Accepted response and returns.
  • Something tied to request A's scope/send (session teardown, or a delayed completion callback from the message router on the same writer) fires again afterward and tries to write http.response.start through the same (now-closed, already-completed) ASGI send, producing the RuntimeError.
  • The reactor appears to mix up these fast-fail paths tightly enough that it can spin instead of cleanly terminating, which is consistent with what we saw during the wedge (sustained CPU, no response to new requests, had to be killed externally).

I don't have a minimal repro yet — this only shows up under organic concurrent load against a long-running stateful session, not in a quick scripted test. Happy to help instrument/reproduce if a maintainer can point at the likely callback path (e.g. is there a completion/cancel callback registered per-request that isn't being cancelled/unregistered when the 202 response path returns early?).

Impact

  • Usually harmless (logged, swallowed, next request succeeds).
  • At least once, coincided with the whole server task group wedging (high CPU, unresponsive, required external restart). Severity assessment is tentative since causation isn't proven, but flagging given the potential for a hang in a long-running server.

What we've done in the meantime

Running a watchdog that health-checks the MCP endpoint and does a bounded, backoff-limited restart (not an unconditional restart loop) with alerting if it doesn't recover after a few attempts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    v1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions