Conversation
The default mempool reserves ~2x device memory of virtual address space on first use, which cannot be satisfied in a 39-bit address space (riscv64 Sv39) and made the test flaky. Create a 2 MiB pool instead and set it before querying it, since the pool getters also create the default pool.
|
/ok to test a8c7196 |
Mirror the Cython change: create a 2 MiB pool and set it before querying it, so the test no longer creates the default pool and no longer needs the xfail_if_mempool_oom guard.
|
/ok to test 870ea5d |
|
rwgk
left a comment
There was a problem hiding this comment.
I reviewed this PR with codex, then went ahead and made the one simple suggested fix under commit ff06d41.
I'm not familiar with the tested production code, so I used codex to narrate the changes with relevant background to me. I'm approving based on what I learned through that narration. (I'm intentionally not posting the narration because it's tailored to me.)
codex GPT-6.1-Sol ultra findings
[P2] Both mempool tests need serialization. The new Python Set/Get sequence and Cython equivalent each install a distinct pool on device 0. CUDA’s current pool is device state, and free-threaded CI runs four concurrent copies. This permits:
- Worker A sets pool A.
- Worker B sets pool B.
- Worker A reads pool B and fails the identity assertion.
Previously, all workers selected the same default pool. Add @pytest.mark.thread_unsafe(reason="changes the device's current memory pool") to both tests; the pytest plugin supports this serialization.
I found no other substantive defects. The capped-pool approach fits the background, handle conversions look correct, and CUDA documents that destroying the current pool restores the default selection. Two nonblocking points remain: Cython cleanup should mirror Python’s try/finally, and removing the default-pool getters deliberately reduces direct API coverage.
I reviewed both changed files at 870ea5d, the full Slack thread through slack-local-bridge, and issue #2381. All 121 checks are successful or skipped. Logs show both affected tests passing in Linux free-threaded CI and Windows MCDM CI; those passes do not exclude the race. RISC-V/Sv39 behavior remains unverified, and the local environment has no accessible NVIDIA driver for reproduction.
|
/ok to test 02fe137 |
Description
follow-on to #2381
test_interop_memPoolin thecuda_bindingsCython and Python tests used the device's default memory pool. Creating the default pool reserves virtual address space of roughly twice the installed device memory, and that reservation is not returned for the lifetime of the process. On address-space-constrained systems this can fail withCUDA_ERROR_OUT_OF_MEMORYeven when plenty of device memory is free, which is the failure mode described in #2381. It also follows the convention incuda_core/tests/AGENTS.md: tests that need a pool should create a capped one.This PR changes both tests to:
maxSize= 2 MiB) withcuMemPoolCreatecudaDeviceSetMemPoolcudaDeviceGetMemPooland assert it is the same handle, then pass it tocuDeviceSetMemPoolThe capped pool is set before it is queried because
cuDeviceGetMemPoolandcudaDeviceGetMemPoolalso create the default pool when no pool has been set. Running the test no longer grows the process virtual size (5 GiB before and after, measured locally on an RTX 6000 Ada).Both tests still check the driver/runtime handle round trip. The Python test no longer needs the
xfail_if_mempool_oomguard. The two*GetDefaultMemPoolcalls are no longer made in either test, socuda_bindingsno longer exercises them directly.Checklist