Repository navigation
Possible memory leak in _cdecimal.c with 'z' format #114563
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Jan 25, 2024 - addedextension-modulesC modules in the Modules dirC modules in the Modules dir
on Jan 25, 2024 Inline quote below. Was there a thread specifically about the memory leak?
Haven't looked at it, but seems to be good news that the mpdecimal library picked up z-format support.
Additionally, distributors of Python-11 and Python-12 are advised to revert
the implementation of the z-format specifier in _decimal.c. It contains a
memory leak for large decimals and does not support the "EG" types.mpdecimal-4.0.0 automatically supports the z-format specifier without patches
to _decimal.c.The following patch cleanly reverts b0b836b
and applies to both Python-11 and Python-12:https://www.bytereef.org/contrib/0001-py12-revert-z-format-specifier.patch
@belm0 in case you're not aware, there's some history here (which I won't get into) that may make the author of that announcement hesitant to interact directly with CPython. If we can verify the memory leak exists, we should fix it in CPython.
Reacted by Mark DickinsonThanks for pointing that out, @JelleZijlstra. I agree we should just try and fix it on our side. I don't know how feasible it is to update to a new mpdecimal and remove our 'z' code, if that will fix it.
I recall
debatingdefending that memory handling code in the PEP PR. I wonder if the former Python contributor (having maintained that integration code) suspected a leak by reading the code, or actually observed a leak.I can reproduce the memory leak on main with this code:
from decimal import Decimal while True: d = Decimal('-.508e+41211') format(d, 'D=-z,.44%')
When I run this on my (macOS / Intel, FWIW) machine, RAM usage (observed with
top) grows at around 80 MB per second. If I remove the 'z' from the format specifier, I no longer see a leak.tracemallocconfirms the leak: if I run:import decimal import tracemalloc tracemalloc.start() repeat = 100 before = tracemalloc.take_snapshot() for _ in range(repeat): d = decimal.Decimal('-.508e+41211') format(d, 'D=-z,.44%') after = tracemalloc.take_snapshot() top_stats = after.compare_to(before, 'lineno') for stat in top_stats[:3]: print(stat)
then the first line of the output is:
/Users/mdickinson/Repositories/python/cpython/test.py:9: size=1697 KiB (+1697 KiB), count=100 (+100), average=17.0 KiBand the size grows linearly with
repeat.@belm0 Can you reproduce the above on your machine, and if so do you have bandwidth to investigate further?
does not support the "EG" types
I don't understand / can't reproduce this: the
EandGtypes will only ever produce a zero output (with whatever sign) for a zero input, and in the zero-input case thezflag appears to be working as expected. E.g., for theEcase:mdickinson@lovelace cpython % ./python.exe Python 3.13.0a3+ (heads/main:a768e12f09, Jan 28 2024, 09:45:39) [Clang 15.0.0 (clang-1500.1.0.2.5)] on darwin Type "help", "copyright", "credits" or "license" for more information. >>> from decimal import Decimal >>> format(Decimal('-0'), '.6E') '-0.000000E+6' >>> format(Decimal('-0'), 'z.6E') '0.000000E+6'
Reacted by John Belmonte- mpdecimal is documented to follow PEP-3101:
https://www.bytereef.org/mpdecimal/doc/libmpdec/assign-convert.html#to-string
- The memory leak for static decimals that are promoted to dynamic ones if the coefficient gets too large is documented here:
https://www.bytereef.org/mpdecimal/doc/libmpdec/memory.html
-
You have to call
mpd_delalso on static decimals (unless they are marked constant). -
The patches in the mail cited in the first message contain a general fallback that catches all specifiers that
_pydecimalsupports. Just apply:
https://www.bytereef.org/contrib/0001-main-revert-z-format-specifier.patch
https://www.bytereef.org/contrib/0002-main-fallback-to-pydecimal-format.patch- There is no need to upgrade to mpdecimal-4.0.0. mpdecimal is a library, distributions ship it and it is trivial to provide a nuget package (like the Windows build does for other libraries).
Finally, @JelleZijlstra, prevailing open source conventions would dictate that upstream is contacted when a new feature is desired in a library, not the other way round.
@mdickinson Thank you for reproducing the memory leak! "EG" was a typo (introduced while coordinating the large amount of patches) I meant "F":
>>> from _decimal import * >>> Decimal("-6.24E-323").__format__("\U000d3219> z3,.10F") '-0.0000000000' >>> >>> from _pydecimal import * >>> Decimal("-6.24E-323").__format__("\U000d3219> z3,.10F") ' 0.0000000000'Reacted by Mark Dickinson@belm0 I did not "maintain" the integration code, I am the sole author of
Modules/_decimal/*, includingmpdecimal.Thank you for the additional links.
prevailing open source conventions would dictate that upstream is contacted when a new feature is desired in a library, not the other way round
I neglected to consider this precisely because distributions may ship/override mpdecimal, so putting an implementation in Python seemed the only way to provide the enhancement for all cases.
@mdickinson what do you think of the fallback approach in the cited patch? It seems like less tricky code to maintain.
>>> Decimal("-6.24E-323").__format__("\U000d3219> z3,.10F") '-0.0000000000'Do we need a separate issue for supporting
F?prevailing open source conventions would dictate that upstream is contacted when a new feature is desired in a library, not the other way round
I neglected to consider this precisely because distributions may ship/override mpdecimal, so putting an implementation in Python seemed the only way to provide the enhancement for all cases.
Yes, I agree, up to a point. For example, all relevant distributions are now
>= 2.5.1:https://repology.org/project/mpdecimal/versions
So now instances of
CONFIG_64in_decimal.ccould be replaced byMPD_CONFIG_64. This is unrelated to this issue, just an example that workaround code can be deleted at some point. It will take a year for them to pick up4.0.0of course!I'm very glad that you mention the difficulties of the coordination issue! I have received very little empathy in 2020 when I tried to fix an incorrect (and unreleased!) patched libmpdec in and for Debian before the freeze.
Do we need a separate issue for supporting
F?No, the fallback code fixes everything including hash formatting. mpdecimal speedups will be picked up automatically when available.
I opened PR #114879 based on Stefan's patches.
what do you think of the fallback approach in the cited patch?
Seems reasonable to me: it makes a lot of sense to me to separate the (somewhat) frequently-changing Python-specific formatting details from the standards-based decimal core. It's a bit ugly to have some of the formatting being done in the C code and some in Python, but that seems like the most pragmatic compromise. (Moving all formatting to Python sounds nice in theory, but would almost certainly have an unacceptable performance impact for common cases.)
Thanks for the PR. I'll review shortly.
- added a commit that references this issue
on Feb 13, 2024
Bug report
Bug description:
See https://mail.python.org/archives/list/python-announce-list@python.org/thread/DHZROL7YYJZTPWJQ3WME4HI3Z65K2H4F/
This feature was originally implemented in #30049
I have not verified the memory leak or looked at Stefan's suggestions.
@mdickinson @belm0
CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs