Skip to content

_remote_debugging: raising Invalid bytes length on codeobjects with linetable > 4096 #151022

Description

@maurycy

Bug report

Bug description:

I've reproduced the issue reported in https://discuss.python.org/t/tachyon-97-error-rate/107619

The issue is caused by:

linetable = read_py_bytes(unwinder,
GET_MEMBER(uintptr_t, code_object, unwinder->debug_offsets.code_object.linetable), 4096);

The linetable for HttpCli.tx_browser and HttpCli.run is above 4096:

2026-06-06T17:13:54.694708000+0200 maurycy@gimel /Users/maurycy/work/cpython (main fded34d*?) % uv run linetable_dist.py ~/Desktop/copyparty
objects=3472 median=51 max=19536 >4096=7
   19536  AuthSrv._reload           /Users/maurycy/Desktop/copyparty/copyparty/authsrv.py:1762
   10640  HttpCli.tx_browser        /Users/maurycy/Desktop/copyparty/copyparty/httpcli.py:6749
    9131  HttpCli.run               /Users/maurycy/Desktop/copyparty/copyparty/httpcli.py:335
    6245  <module>                  /Users/maurycy/Desktop/copyparty/copyparty/util.py:1
    5087  Up2k._handle_json         /Users/maurycy/Desktop/copyparty/copyparty/up2k.py:3035
    4146  SvcHub.__init__           /Users/maurycy/Desktop/copyparty/copyparty/svchub.py:136
    4133  HttpCli.handle_plain_upload  /Users/maurycy/Desktop/copyparty/copyparty/httpcli.py:3648
    3753  HttpCli.dump_to_file      /Users/maurycy/Desktop/copyparty/copyparty/httpcli.py:2444

The 4096 limit is too small even for our stdlib:

2026-06-06T17:14:47.179006000+0200 maurycy@gimel /Users/maurycy/work/cpython (main fded34d*?) % uv run linetable_dist.py ~/work/cpython/Lib 
objects=92035 median=49 max=37416 >4096=37
   37416  <module>                  /Users/maurycy/work/cpython/Lib/html/entities.py:1
   20958  <module>                  /Users/maurycy/work/cpython/Lib/test/test_ast/snippets.py:1
   17060  <module>                  /Users/maurycy/work/cpython/Lib/locale.py:1
   11158  <module>                  /Users/maurycy/work/cpython/Lib/test/test_dis.py:1
   10063  SuggestionFormattingTestBase.test_name_error_suggestions_do_not_trigger_for_too_many_locals.<locals>.func  /Users/maurycy/work/cpython/Lib/test/test_traceback.py:4899
    9953  <module>                  /Users/maurycy/work/cpython/Lib/stringprep.py:1
    9400  <module>                  /Users/maurycy/work/cpython/Lib/test/re_tests.py:1
    5892  <module>                  /Users/maurycy/work/cpython/Lib/encodings/aliases.py:1
    5791  <module>                  /Users/maurycy/work/cpython/Lib/test/test_socket.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/mac_arabic.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/cp866.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/cp865.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/cp863.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/cp862.py:1
    5730  <module>                  /Users/maurycy/work/cpython/Lib/encodings/cp861.py:1

Reproduction

Using real file from stdlib:

import os, _remote_debugging

src = open("/Users/maurycy/work/cpython/Lib/html/entities.py").read()
src += "\n_remote_debugging.RemoteUnwinder(os.getpid()).get_stack_trace()\n"
exec(compile(src, "entities.py", "exec"))
print("OK")

We get:

2026-06-06T17:17:34.919627000+0200 maurycy@gimel /Users/maurycy/work/cpython (main fded34d*?) % uv run repro.py
Traceback (most recent call last):
  File "/Users/maurycy/src/github.com/maurycy/cpython/repro.py", line 5, in <module>
    exec(compile(src, "entities.py", "exec"))
    ~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "entities.py", line 2515, in <module>
RuntimeError: Invalid bytes length (37445) at 0xc5091c000

Bonus

(thx Claude)

linetable_dist.py
#!/usr/bin/env python3
import warnings
from pathlib import Path
from statistics import median
from sys import argv
from types import CodeType

warnings.simplefilter("ignore")
SKIP = {"venv", "site-packages", "__pycache__"}


def codes(co):
    yield co
    for c in co.co_consts:
        if isinstance(c, CodeType):
            yield from codes(c)


rows = []
for arg in argv[1:]:
    p = Path(arg).expanduser()
    for f in [p] if p.is_file() else p.rglob("*.py"):
        if SKIP & set(f.parts):
            continue
        try:
            co = compile(f.read_text(errors="replace"), str(f), "exec")
        except Exception:
            continue
        rows += [(len(c.co_linetable), str(f), c.co_qualname, c.co_firstlineno)
                 for c in codes(co)]

rows.sort(reverse=True)
vals = [s for s, *_ in rows]

print(f"objects={len(vals)} median={median(vals):.0f} max={max(vals)} "
      f">4096={sum(v > 4096 for v in vals)}")
for s, fn, qual, lineno in rows[:15]:
    print(f"{s:8d}  {qual:24s}  {fn}:{lineno}")

CPython versions tested on:

CPython main branch

Operating systems tested on:

macOS

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    on Jun 6, 2026
  2. maurycy commented on Jun 6, 2026

    @maurycy
    ContributorAuthor

    cc @pablogsal 2^16 is tempting :-)

  3. pablogsal commented on Jun 6, 2026

    @pablogsal
    Member

    640K ought to be enough for anybody

  4. goutamadwant commented on Jun 6, 2026

    @goutamadwant
    Contributor

    @maurycy I could reproduce this on current main with the html.entities example. The line table is 37445 bytes, so the failure happens before parsing because the remote bytes read is capped at 4096.

    I opened #151036 with a narrow fix. It keeps the generic remote bytes reader bounded, but uses a separate larger bound for code object line tables. I also added a regression test that builds a code object with a line table over 4096 bytes and verifies that same-process unwinding succeeds.

  5. added a commit that references this issue on Jun 6, 2026
  6. maurycy commented on Jun 9, 2026

    @maurycy
    ContributorAuthor

    @pablogsal The more I think about this... For the purpose of mach_vm_remap, I'm trying to measure the bias (ie: delta between the truth and what we report, with aliasing and sampling rates between 2-5MHz I started to wonder about it; spoiler: we're great here.) But the original issue reports heavy bias actually... linetable, qualname/filename and MAX_FRAMES (and this _sync_coordinator.py; unlucky person who accidentially uses this name too) drop the whole frame silently. Now, for a long-running profiler on a large-scale production system a heavy bias, like the one observed originally, makes the profiling not trustworthy, and time completely wasted. In other words, apart from making the limits realistic, I believe we should not drop the frames on limit but degrade

  7. pablogsal commented on Jun 9, 2026

    @pablogsal
    Member

    @pablogsal The more I think about this... For the purpose of mach_vm_remap, I'm trying to measure the bias (ie: delta between the truth and what we report, with aliasing and sampling rates between 2-5MHz I started to wonder about it; spoiler: we're great here.) But the original issue reports heavy bias actually... linetable, qualname/filename and MAX_FRAMES (and this _sync_coordinator.py; unlucky person who accidentially uses this name too) drop the whole frame silently. Now, for a long-running profiler on a large-scale production system a heavy bias, like the one observed originally, makes the profiling not trustworthy, and time completely wasted. In other words, apart from making the limits realistic, I believe we should not drop the frames on limit but degrade

    What kind of degradation did you have in mind? Emitting the frame with line unknown when the linetable read fails, or something more specific?

  8. maurycy commented on Jun 9, 2026

    @maurycy
    ContributorAuthor

    @pablogsal The more I think about this... For the purpose of mach_vm_remap, I'm trying to measure the bias (ie: delta between the truth and what we report, with aliasing and sampling rates between 2-5MHz I started to wonder about it; spoiler: we're great here.) But the original issue reports heavy bias actually... linetable, qualname/filename and MAX_FRAMES (and this _sync_coordinator.py; unlucky person who accidentially uses this name too) drop the whole frame silently. Now, for a long-running profiler on a large-scale production system a heavy bias, like the one observed originally, makes the profiling not trustworthy, and time completely wasted. In other words, apart from making the limits realistic, I believe we should not drop the frames on limit but degrade

    What kind of degradation did you have in mind? Emitting the frame with line unknown when the linetable read fails, or something more specific?

    Yes, only this, truncation or a placeholder in case of qualname/filename or MAX_FRAMES etc. I think. More than open for any other ideas.

  9. added 2 commits that reference this issue on Jun 11, 2026
  10. chris-eibl commented on Jul 18, 2026

    @chris-eibl
    Member

    FWIW, when profiling pystone.py1,

    python.exe -m profiling.sampling run pystone.py 100000
    

    I also get a high error rate (about 30%) on Windows (if that makes a difference).

    The attached PR doesn't help, and also from the errors I think this must be something different:
    Lots of
    Failed to parse initial frame in chain
    ReadProcessMemory failed for PID 24180 at address 0xc0000000 (size 64, partial read 0 bytes): Windows error 998
    ReadProcessMemory failed for PID 24180 at address 0xe7 (size 64, partial read 0 bytes): Windows error 299

    Where ERROR_NOACCESS=998 and ERROR_PARTIAL_COPY=299, which all indicate reading from invalid memory (already quite obvious from the pointer values).

    Footnotes

    1. e.g. https://github.com/dundee/pybenchmarks/blob/master/bencher/programs/pystone/pystone.python3 ↩

  11. added a commit that references this issue on Jul 18, 2026
  12. goutamadwant commented on Jul 18, 2026

    @goutamadwant
    Contributor

    @chris-eibl Thanks for testing this on Windows. I updated the PR to use a 64 KiB line-table limit and removed the separate entry-count limit.

    The ReadProcessMemory errors shown here occur while resolving the initial frame chain, before line-table parsing begins, so they appear to be a separate failure path from the large-line-table issue addressed by #151036.

    Could you compare the same pystone run with --blocking? If the errors disappear, that would help determine whether the invalid frame addresses come from a race in the non-blocking snapshot path. I think the Windows frame-chain failure should otherwise be tracked separately so both problems remain independently reproducible. Let me know.
    cc @maurycy

  13. chris-eibl commented on Jul 18, 2026

    @chris-eibl
    Member

    Yeah, --blocking drastically reduces the error rate (after applying #152471). And I agree this is a different (and most probably only Windows related) issue.

  14. chris-eibl commented on Jul 19, 2026

    @chris-eibl
    Member

    Having a closer look, I think this quote from the documentation

    However, non-blocking sampling can occasionally produce incomplete or inconsistent stack traces [...] in programs with very fast-changing call stacks where functions enter and exit between the start and end of a single stack read

    very much applies to pystone.py. I've just wondered about the high error rate without giving it a closer thought. Sorry for the noise here ...

  15. sergey-miryanov commented on Jul 19, 2026

    @sergey-miryanov
    Contributor

    Some background from py-spy but the reason the same, IIUC: https://www.benfrederickson.com/why-python-needs-paused-during-profiling/

  16. added a commit that references this issue on Jul 19, 2026
  17. added a commit that references this issue on Jul 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions