Skip to content

Slow scan: uv/pylock hosted 4.1x/3.0x median ms/pkg (per-patch lock re-serialize) #836

Description

[agent] Bench: uv and pylock (PEP 751 pylock.toml) have both been flagged slow in 4 consecutive benchmark runs (2026-10-02 through 10-05). Their hosted scans cost about 3–4x the median ms per package across all package managers. They share one code path (PythonLockSession), so this issue covers both. Filing was deferred from the 10-04 run by the 3-new-issues-per-run cap.

run uv/hosted uv ms/pkg (x median) pylock/hosted pylock ms/pkg (x median)
2026-10-02 136.0 ms 0.340 (flagged) 113.6 ms 0.284 (flagged)
2026-10-03 143.9 ms 0.360 (2.8x) 142.1 ms 0.355 (2.8x)
2026-10-04 127.1 ms 0.318 (2.8x) 126.2 ms 0.316 (2.8x)
2026-10-05 244.0 ms 0.610 (4.1x) 179.8 ms 0.449 (3.0x)

Fixtures: 400 packages with 12 patched, so ms/pkg is high even at modest wall times. 10-05 ran on a slower runner, where the absolute times are about 1.4x those of earlier runs. The ratio to the cross-PM median is the signal.

Hot spot (callgrind, uv/hosted and pylock/hosted, 2026-10-05, main 045d7ec7)

  • utils::python_lock::PythonLockSession::rewrite: 39% (uv) / 38% (pylock) of all instructions. Most of it is toml_edit Display for DocumentMut (31% / 35% inclusive, under encode::visit_table).
  • Cause: the lock is parsed once per session, but PythonLockSession::rewrite ends with document.to_string(). The caller (patch/redirect/mod.rs ~L4675, the for &(dep, sha256) in &usable loop) calls it once per patched dep and then diffs the full old and new text in record_python_lock_edits. So serialization is O(patches × lock size).
  • Possible fix: apply every planned edit to the one DocumentMut, serialize once after the loop, and compute per-dep edit records from the document (or one final diff) instead of from full-text renderings.
  • Secondary: vex::discover::pypi_locks::extract is 23% (uv) / 31% (pylock). Wiring discovery parses the same lock again with toml_edit (toml_or_diag), separately from the redirect session. Sharing the parsed document would save another full parse.

Same shape as #760 (poetry) and #762 (pdm), which re-parse per patch; here it's the re-serialization.

Runner

4 vCPU, Intel(R) Xeon(R) Processor @ 2.80GHz (cloud sandbox), main 045d7ec7. The weekly A/B vs 2463257a (#277) is flat (uv/hosted −0.5%, pylock/hosted +3.8%), so this is standing cost, not a regression.

Repro

CARGO_PROFILE_PERF_INHERITS=release CARGO_PROFILE_PERF_LTO=thin CARGO_PROFILE_PERF_STRIP=none \
  cargo build --locked --profile perf -p socket-patch-cli -p socket-patch-bench
target/perf/socket-patch-bench run --bin target/perf/socket-patch -f '^uv/' -f '^pylock/' -v
# profile: serve the fixture, then prefix the printed command's binary with
#   valgrind --tool=callgrind
target/perf/socket-patch-bench serve uv/hosted --bin target/perf/socket-patch

Tracked in the ledger of #575 (standing slow-systems list).


Generated by Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions