Skip to content

broken link to A.Neumaier article in built-in sum comment #111933

Description

@enzbus

Bug report

Bug description:

The new implementation of sum on Python 3.12 (cfr. #100425 , #100426 , #107785 ) is not associative on simple input values. This minimal code shows the bug:

On Python 3.11:

>>> a = [0.1, -0.2, 0.3, -0.4, 0.5]
>>> a.append(-sum(a))
>>> sum(a) == 0
True

On Python 3.12:

>>> a = [0.1, -0.2, 0.3, -0.4, 0.5]
>>> a.append(-sum(a))
>>> sum(a) == 0
False

I'm sure this affects more users than the "improved numerical accuracy" on badly scaled input data which most users don't ever deal with, and for which exact arithmetic is already available in the Standard Library
-> https://docs.python.org/3/library/decimal.html.

I'm surprised this low-level change was accepted with so little review. There are other red flags connected with this change:

Is anybody interested in keeping the quality of cPython's codebase high? When I learned Python, I remember one of the first thing in the official tutorial was that Python is a handy calculator, and now to me it seems broken. @gvanrossum ?

CPython versions tested on:

3.12

Operating systems tested on:

Linux

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    on Nov 10, 2023
  2. skirpichev commented on Nov 10, 2023

    @skirpichev
    Member

    The researcher to which the algorithm is credited has an empty academic page, with no PDFs -> https://www.mat.univie.ac.at/~neum/

    FYI, citation: Neumaier, A. (1974), Rundungsfehleranalyse einiger Verfahren zur Summation endlicher Summen. Z. angew. Math. Mech., 54: 39-51. https://doi.org/10.1002/zamm.19740540106.

    Full text is still available in the archive: https://web.archive.org/web/20220804051351/https://www.mat.univie.ac.at/~neum/scan/01.pdf Should we use this url or above citation (no free text access)?

    BTW, I doubt there is an issue, besides a broken link. Associativity is broken for addition of floating point numbers, see e.g. TAOCP, Vol. 2, 4.2.2. The Tutorial has a dedicated section for floating point arithmetic, which cover this "issue".

    exact arithmetic is already available in the Standard Library -> https://docs.python.org/3/library/decimal.html

    The decimal module actually not about exact arithmetics, rather the fractions module.

  3. Eclips4 commented on Nov 10, 2023

    @Eclips4
    Member
  4. added
    3.12only security fixes
    3.13only security fixes
    on Nov 10, 2023
  5. added a commit that references this issue on Nov 10, 2023
  6. terryjreedy commented on Nov 10, 2023

    @terryjreedy
    Member

    One might naively expect sum(a) to be .9 -.6 = .3. But it is actually 0.29999999999999993, so the error is about 7e-17. The sum(a2) is about 3e-17, so the absolute error has gone down. 0.0 error is something of an accident.

    Now add -.3 to list a and sum. Instead of the naively expected 0.0, I got -5.6e-17 in 3.11 and -2.8e-17 in 3.12 (both rounded to 2 digits), which is half the error. This is at least as realistic example than the original one and has nothing to do with badly scaled data.

    @enzbus In the future, please first ask about possible bugs on help forums such as https://discuss.python.org/c/users/7. And please omit disparaging comments wherever you post.

  7. sweeneyde commented on Nov 10, 2023

    @sweeneyde
    Member

    You also don't have to modify the example too much to see a case where the accuracy of the sum function improved with the change:

    >>> # Python 3.11
    >>> sum([0.1, -0.2, 0.3, -0.4, 0.5, -0.1, 0.2, -0.3, 0.4, -0.5])
    -5.551115123125783e-17
    
    >>> # Python 3.12
    >>> sum([0.1, -0.2, 0.3, -0.4, 0.5, -0.1, 0.2, -0.3, 0.4, -0.5])
    0.0
    
  8. enzbus commented on Nov 10, 2023

    @enzbus
    Author

    I'm actually getting the error on random inputs, see the commit above which references this bug report. The point remains, this is a regression, adds little value (as I explain above), is the implementation of a 1974 paper in german (really?) which is not online any more, and should not have happened.

  9. enzbus commented on Nov 10, 2023

    @enzbus
    Author

    @terryjreedy please re-open, at least to see if this affects other users. Please note that there were no disparaging comments. I noted that the Standard Library already provides arbitrary precision arithmetic in the decimal module, this change refers a suspicious article in German which I couldn't even find (and I went on the academic webpage of the researcher). I think this should not have passed quality review, since this bug is real (I provided a minimal code to show it, but I encountered it on random inputs where the comparison is exact on python < 3.12), especially since it adds little value.

  10. skirpichev commented on Nov 10, 2023

    @skirpichev
    Member

    @terryjreedy, lets not forgot the real issue with the broken link, see PR #111937.

  11. terryjreedy commented on Nov 10, 2023

    @terryjreedy
    Member

    Everyone who uses non-integral binary floats in any language, not just Python, is affected by the quirks of decimal-binary conversions and of floating-point arithmetic. Comparing the results with == is usually a bug. One solution is to choose a unit such that all values are integral.

    @enzbus Suggesting that most or all of us core developers do not care about code quality is not a compliment.

    @skirpichev You are right. Reopening for the doc fix.

  12. 17 remaining items

  13. pochmann4 commented on Nov 11, 2023

    @pochmann4

    @kirpichevs Yes, I used decimal as well. I used precision 99, though. Not sure what to think of your calculations using the default precision 28, which is not enough here for exact calculations. (Might be enough for your point, I'm just not sure. But with 99, I get sd = 2.7755575615628914e-17 (in Python 3.11, but I think that's the same in 3.12 and main).)

  14. kirpichevs commented on Nov 11, 2023

    @kirpichevs

    I used precision 99, though. Not sure what to think of your calculations using the default precision 28, which is not enough here for exact calculations.

    Well, with this precision there is still a difference with sum() v3.12. But with prec=99:

    >>> getcontext().prec=99
    >>> a = [0.1, -0.2, 0.3, -0.4, 0.5]
    >>> ad = list(map(Decimal, a))
    >>> Decimal(float(reduce(add, ad)))
    Decimal('0.29999999999999993338661852249060757458209991455078125')
    >>> float(reduce(add, ad + [-Decimal(float(reduce(add, ad)))]))
    2.7755575615628914e-17
    >>> sd = _
    >>> sum(a + [-sum(a)])
    2.7755575615628914e-17
    >>> _ - sd
    0.0
  15. enzbus commented on Nov 11, 2023

    @enzbus
    Author

    The exact values stored in a are:

    ...

    And the exact sum is

    In fact, we have decimal module to demonstrate your argument (current main branch):

    >>> from decimal import Decimal
    
    >>> from functools import reduce
    
    >>> from operator import add
    
    >>> a = [0.1, -0.2, 0.3, -0.4, 0.5]
    
    >>> ad = list(map(Decimal, a))
    
    >>> Decimal(float(reduce(add, ad)))
    
    Decimal('0.29999999999999993338661852249060757458209991455078125')
    
    >>> float(reduce(add, ad + [-Decimal(float(reduce(add, ad)))]))
    
    2.77555756156e-17
    
    >>> sd = _
    
    >>> sum(a + [-sum(a)])
    
    2.7755575615628914e-17
    
    >>> sr = _
    
    >>> abs((sr - sd)/(0.0 - sd))
    
    1.0417222640068697e-12
    

    In fact, new sum() here more accurate than naive summation for several orders of magnitude.

    (And please stop blocking me.)

    @pochmann, you are not alone. (Is he already blocking core developers or not yet?)

    If an == comparison returns correctly True

    @enzbus, people already told you, that using == to compare floating point numbers is a bug (in your program, yes). E.g. floating point arithmetic lacks associativity (for addition or multiplication), so in general the result will depend on the order of evaluation. Your == comparison will be fragile just for that reason.

    You are really trying to shift blame here, away from having introduced a breaking change. If you don't like my code, that's fine, but I test it across all supported CPython versions, and when I added 3.12 I had to go and find why an expression that was exact became inexact. It took me a while because I would never have thought CPython was to blame. But now I understand, I think I see a cavalier attitude among some core devs on introducing breaking changes.

  16. ericvsmith commented on Nov 11, 2023

    @ericvsmith
    Member

    I don't think anyone is denying it's a breaking change. It clearly was.

    However, it's a breaking change we were comfortable making in a effort to improve Python. Every release of CPython has breaking changes, in at least the sense that the changes can be detected somehow. Which is unfortunate, but such is the cost of progress. I'm sorry you were affected by this one. But I agree with others that relying on == with floats is a fragile test, and should be avoided.

    There's no appetite here for reverting this change.

  17. kirguesswho commented on Nov 11, 2023

    @kirguesswho

    why an expression that was exact became inexact.

    Could you, please, explain (you can still block people here, enjoy!) why do you think it "was exact"?

    Consider a simple list of floats:

    >>> a = [0.1, -0.2, 0.3, -0.4, 0.2]

    Now we try to compute sum:

    >>> from random import sample
    >>> from functools import reduce
    >>> from operator import add
    >>> reduce(add, sample(a, 5))
    0.0
    >>> reduce(add, sample(a, 5))
    -8.326672684688674e-17
    >>> reduce(add, sample(a, 5))
    0.0
    >>> reduce(add, sample(a, 5))
    -2.7755575615628914e-17

    Oops. Different results! Which answer do you prefer and why?

    PS:

    since you mentioned, #100425 is dated December 2022, and the researcher from Vienna says he moved out his website in April 2022 (on the bottom, since people also took issue with the wording of his academic webpage). So the link was dead

    As it was pointed from the beginning of the thread, the link was not dead at least till August 2022. Your argument is wrong.

  18. enzbus commented on Nov 11, 2023

    @enzbus
    Author

    I don't think anyone is denying it's a breaking change. It clearly was.

    However, it's a breaking change we were comfortable making in a effort to improve Python. Every release of CPython has breaking changes, in at least the sense that the changes can be detected somehow. Which is unfortunate, but such is the cost of progress. I'm sorry you were affected by this one. But I agree with others that relying on == with floats is a fragile test, and should be avoided.

    There's no appetite here for reverting this change.

    Ok, apologies accepted. Thanks for making CPython available to the community.

  19. Eclips4 commented on Nov 11, 2023

    @Eclips4
    Member

    Let's make it's open until PR with fix(inaccessible link to the research) get merged.

  20. locked as too heated and limited conversation to collaborators on Nov 11, 2023
  21. terryjreedy commented on Nov 11, 2023

    @terryjreedy
    Member

    We are not going to revert the change to sum. We are going to improve the reference. No more discussion needed.

  22. rhettinger commented on Nov 12, 2023

    @rhettinger
    Contributor

    Let's make it's open until PR with fix(inaccessible link to the research) get merged.

    That's now done.

  23. changed the title [-]New CPython 3.12 implementation of `sum` breaks compatibility; it inserts errors on expressions that were correct on all previous CPython versions[/-] [+]Docs: Invalid link to the research paper on the new `sum` implementation[/+] on Sep 26, 2024
  24. changed the title [-]Docs: Invalid link to the research paper on the new `sum` implementation[/-] [+]broken link to A.Neumaier article in built-in sum comment[/+] on Sep 26, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.12only security fixes3.13only security fixes

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions