Skip to content

Quadratic time internal base conversions #90716

Description

@tim-one
BPO 46558
Nosy @tim-one, @cfbolz, @sweeneyde
Files
  • todecstr.py
  • todecstr.py
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = <Date 2022-01-28.03:12:39.839>
    created_at = <Date 2022-01-28.02:31:44.271>
    labels = ['interpreter-core', 'performance']
    title = 'Quadratic time internal base conversions'
    updated_at = <Date 2022-01-31.06:11:48.110>
    user = 'https://github.com/tim-one'

    bugs.python.org fields:

    activity = <Date 2022-01-31.06:11:48.110>
    actor = 'tim.peters'
    assignee = 'none'
    closed = True
    closed_date = <Date 2022-01-28.03:12:39.839>
    closer = 'tim.peters'
    components = ['Interpreter Core']
    creation = <Date 2022-01-28.02:31:44.271>
    creator = 'tim.peters'
    dependencies = []
    files = ['50593', '50595']
    hgrepos = []
    issue_num = 46558
    keywords = []
    message_count = 9.0
    messages = ['411962', '411966', '411969', '411971', '412120', '412122', '412172', '412191', '412192']
    nosy_count = 3.0
    nosy_names = ['tim.peters', 'Carl.Friedrich.Bolz', 'Dennis Sweeney']
    pr_nums = []
    priority = 'normal'
    resolution = 'wont fix'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'performance'
    url = 'https://bugs.python.org/issue46558'
    versions = []

    Activity

    1. tim-one commented on Jan 28, 2022

      @tim-one
      MemberAuthor

      Our internal base conversion algorithms between power-of-2 and non-power-of-2 bases are quadratic time, and that's been annoying forever ;-) This applies to int<->str and int<->decimal.Decimal conversions. Sometimes the conversion is implicit, like when comparing an int to a Decimal.

      For example:

      >> a = 1 << 1000000000 # yup! a billion and one bits
      >> s = str(a)

      I gave up after waiting for over 8 hours, and the computation apparently can't be interrupted.

      In contrast, using the function in the attached todecstr.py gets the result in under a minute:

      >>> a = 1 << 1000000000
      >>> s = todecstr(a)
      >>> len(s)
      301029996

      That builds an equal decimal.Decimal in a "clever" recursive way, and then just applies str to _that_.

      That's actually a best case for the function, which gets major benefit from the mountains of trailing 0 bits. A worst case is all 1-bits, but that still finishes in under 5 minutes:

      >>> a = 1 << 1000000000
      >>> s2 = todecstr(a - 1)
      >>> len(s2)
      301029996
      >>> s[-10:], s2[-10:]
      ('1787109376', '1787109375')

      A similar kind of function could certainly be written to convert from Decimal to int much faster, but it would probably be less effective. These things avoid explicit division entirely, but fat multiplies are key, and Decimal implements a fancier * algorithm than Karatsuba.

      Not for the faint of heart ;-)

    2. sweeneyde commented on Jan 28, 2022

      @sweeneyde
      Member
    3. tim-one commented on Jan 28, 2022

      @tim-one
      MemberAuthor

      Dennis, partly, although that was more aimed at speeding division, while the approach here doesn't use division at all.

      However, thinking about it, the implementation I attached doesn't actually for many cases (it doesn't build as much of the power tree in advance as may be needed). Which I missed because all the test cases I tried had mountains of trailing 0 or 1 bits, not mixtures.

      So I'm closing this anyway, at least until I can dream up an approach that always works. Thanks!

    4. tim-one commented on Jan 28, 2022

      @tim-one
      MemberAuthor

      Changed the code so that inner() only references one of the O(log log n) powers of 2 we actually precomputed (it could get lost before if lo was non-zero but within n had at least one leading zero bit - now we pass the conceptual width instead of computing it on the fly).

    5. tim-one commented on Jan 30, 2022

      @tim-one
      MemberAuthor

      The test case here is a = (1 << 100000000) - 1, a solid string of 100 million 1 bits. The goal is to convert to a decimal string.

      Methods:

      native: str(a)

      numeral: the Python numeral() function from bpo-3451's div.py after adapting to use the Python divmod_fast() from the same report's fast_div.py.

      todecstr: from the Python file attached to this report.

      gmp: str() applied to gmpy2.mpz(a).

      Timings:

      native: don't know; gave up after waiting over 2 1/2 hours.
      numeral: about 5 1/2 minutes.
      todecstr: under 30 seconds. (*)
      gmp: under 6 seconds.

      So there's room for improvement ;-)

      But here's the thing: I've lost count of how many times someone has whipped up a pure-Python implementation of a bigint algorithm that leaves CPython in the dust. And they're generally pretty easy in Python.

      But then they die there, because converting to C is soul-crushing, losing the beauty and elegance and compactness to mountains of low-level details of memory-management, refcounting, and checking for errors after every tiny operation.

      So a new question in this endless dilemma: _why_ do we need to convert to C? Why not leave the extreme cases to far-easier to write and maintain Python code? When we're cutting runtime from hours down to minutes, we're focusing on entirely the wrong end to not settle for 2 minutes because it may be theoretically possible to cut that to 1 minute by resorting to C.

      (*) I hope this algorithm tickles you by defying expectations ;-) It essentially stands numeral() on its head by splitting the input by a power of 2 instead of by a power of 10. As a result no divisions are used. But instead of shifting decimal digits into place, it has to multiply the high-end pieces by powers of 2. That seems insane on the face of it, but hard to argue with the clock ;-) The "tricks" here are that the O(log log n) powers of 2 needed can be computed efficiently in advance of any splitting, and that all the heavy arithmetic is done in the decimal module, which implements fancier-than-Karatsuba multiplication and whose values can be converted to decimal strings very quickly.

    6. tim-one commented on Jan 30, 2022

      @tim-one
      MemberAuthor

      Addendum: the "native" time (for built in str(a)) in the msg above turned out to be over 3 hours and 50 minutes.

    7. cfbolz commented on Jan 30, 2022

      cfbolzmannequin
      Mannequin

      Somebody pointed me to V8's implementation of str(bigint) today:

      https://github.com/v8/v8/blob/main/src/bigint/tostring.cc

      They say that they can compute str(factorial(1_000_000)) (which is 5.5 million decimal digits) in 1.5s:

      https://twitter.com/JakobKummerow/status/1487872478076620800

      As far as I understand the code (I suck at C++) they recursively split the bigint into halves using % 10^n at each recursion step, but pre-compute and cache the divisors' inverses.

    8. tim-one commented on Jan 31, 2022

      @tim-one
      MemberAuthor

      The factorial of a million is much smaller than the case I was looking at. Here are rough timings on my box, for computing the decimal string from the bigint (and, yes, they all return the same string):

      native: 475 seconds (about 8 minutes)
      numeral: 22.3 seconds
      todecstr: 4.10 seconds
      gmp: 0.74 seconds

      "They recursively split the bigint into halves using % 10^n at each recursion step". That's the standard trick for "output" conversions. Beyond that, there are different ways to try to use "fat" multiplications instead of division. The recursive splitting all on its own can help, but dramatic speedups need dramatically faster multiplication.

      todecstr treats it as an "input" conversion instead, using the decimal module to work mostly in base 10. That, by itself, reduces the role of division (to none at all in the Python code), and decimal has a more advanced multiplication algorithm than CPython's bigints have.

    9. tim-one commented on Jan 31, 2022

      @tim-one
      MemberAuthor

      todecstr treats it as an "input" conversion instead, ...

      Worth pointing this out since it doesn't seem widely known: "input" base conversions are _generally_ faster than "output" ones. Working in the destination base (or a power of it) is generally simpler.

      In the math.factorial(1000000) example, it takes CPython more than 3x longer for str() to convert it to base 10 than for int() to reconstruct the bigint from that string. Not an O() thing (they're both quadratic time in CPython today).

    10. 80 remaining items

    11. byeongkeunahn commented on Sep 13, 2023

      @byeongkeunahn

      It seems the num-bigint crate attaches to Python 3 quite well with PyO3. The following image shows the performance of the new multiplication implementation on Ryzen 7 2700X, averaged over 10 runs. For small numbers, the num-bigint crate uses naive, Karatsuba, and Toom-3 before switching to number-theoretic transform. The cross-over point over the CPython native multiplication (version 3.11.3) is around 3,000 bits.

      benchmark-PyO3

      The binding I used is very simple. Further optimizations may be possible.

      use pyo3::prelude::*;
      use num_bigint::BigInt;
      
      #[pyfunction]
      fn mul(x: BigInt, y: BigInt) -> BigInt {
          &x * &y
      }
      
      #[pymodule]
      fn num_bigint_pyo3(_py: Python, m: &PyModule) -> PyResult<()> {
          m.add_function(wrap_pyfunction!(mul, m)?)?;
          Ok(())
      }
      import num_bigint_pyo3
      a, b = 10**100000, 9**100000
      x = num_bigint_pyo3.mul(a, b)
      assert x == a*b
    12. byeongkeunahn commented on Sep 13, 2023

      @byeongkeunahn

      Here are the str -> int conversion timings with the new multiplication routine.

      • 40,000 digits:
      int                           : 0.006869841 ± 0.000509983 sec
      str_to_int_using_decimal      : 0.026808119 ± 0.000423398 sec
      str_to_int                    : 0.003381038 ± 0.000521641 sec
      str_to_int_new_mul            : 0.001501775 ± 0.000504214 sec
      
      • 400,000 digits:
      int                           : 0.678053117 ± 0.005104713 sec
      str_to_int_using_decimal      : 0.553185940 ± 0.003627799 sec
      str_to_int                    : 0.127861691 ± 0.000618099 sec
      str_to_int_new_mul            : 0.023358464 ± 0.000615376 sec
      
      • 4,000,000 digits:
      int                           : 68.152991366 ± 0.353029732 sec
      str_to_int_using_decimal      : 9.126699829 ± 0.067300091 sec
      str_to_int                    : 4.851845384 ± 0.030668179 sec
      str_to_int_new_mul            : 0.326830316 ± 0.002546696 sec
      
      Code
      import sys
      sys.set_int_max_str_digits(0)
      
      import num_bigint_pyo3
      from decimal import *
      from time import time
      from random import choices
      import numpy as np
      
      setcontext(Context(prec=MAX_PREC, Emax=MAX_EMAX, Emin=MIN_EMIN))
      
      pow5 = [5]
      while len(pow5) <= 23:
          pow5.append(num_bigint_pyo3.mul(pow5[-1], pow5[-1]))
      
      def str_to_int_new_mul(s):
          def _str_to_int(l, r):
              if r - l <= 3000:
                  return int(s[l:r])
              lg_split = (r - l - 1).bit_length() - 1
              split = 1 << lg_split
              return (num_bigint_pyo3.mul(_str_to_int(l, r - split), pow5[lg_split]) << split) + _str_to_int(r - split, r)
          return _str_to_int(0, len(s))
      
      def str_to_int(s):
          def _str_to_int(l, r):
              if r - l <= 3000:
                  return int(s[l:r])
              lg_split = (r - l - 1).bit_length() - 1
              split = 1 << lg_split
              return ((_str_to_int(l, r - split) * pow5[lg_split]) << split) + _str_to_int(r - split, r)
          return _str_to_int(0, len(s))
      
      def str_to_int_using_decimal(s, bits=1024):
          d = Decimal(s)
          div = Decimal(2) ** bits
          divs = []
          while div <= d:
              divs.append(div)
              div *= div
          digits = [d]
          for div in reversed(divs):
              digits = [x
                        for digit in digits
                        for x in divmod(digit, div)]
              if not digits[0]:
                  del digits[0]
          b = b''.join(int(digit).to_bytes(bits//8, 'big')
                       for digit in digits)
          return int.from_bytes(b, 'big')
      
      # Benchmark against int() on random 1-million digits string
      funcs = [int, str_to_int_using_decimal, str_to_int, str_to_int_new_mul]
      s = ''.join(choices('123456789', k=1_000_000))
      expect = None
      for f in funcs:
          t_list = []
          for _ in range(10):
              t = time()
              result = f(s)
              t_elapsed = time() - t # f.__name__)
              if expect is None:
                  expect = result
              assert result == expect
              t_list.append(t_elapsed)
          mean, std = np.mean(t_list), np.std(t_list)
          print("{0}: {1:.9f} ± {2:.9f} sec".format(f.__name__.ljust(30), mean, std))
    13. oscarbenjamin commented on Sep 27, 2023

      @oscarbenjamin
      Contributor

      Here are the str -> int conversion timings with the new multiplication routine.

      It looks very nice to me. The question is what you would propose to do going on from here. If the suggestion is that CPython might depend on this (even in an optional way) then there are many details that would need to be worked out before that could be considered.

    14. byeongkeunahn commented on Sep 29, 2023

      @byeongkeunahn

      It looks very nice to me. The question is what you would propose to do going on from here. If the suggestion is that CPython might depend on this (even in an optional way) then there are many details that would need to be worked out before that could be considered.

      Yes, I'd love to see CPython's multiplication improved asymptotically. It would also be great to improve other arithmetic operations too, but as a first step, I think it would be appropriate to limit the scope to multiplication.

      Currently the binding through PyO3 receives the raw bits of the integer through _PyLong_ToByteArray and sends it back to Python via _PyLong_FromByteArray. I think that mechanism can still work in the C code, which might help lessen the maintenance burden since the existing APIs don't need to be modified substantially.

      On the other hand, more deliberation is needed to port the primitive operations from Rust to C. Rust provides portable primitives such as add-with-carry and 128bit integer. Although it is possible to access the higher 64bit of a 64 x 64 multiplication in most C compilers, this would not be portable. Switching to plain portable C (without compiler intrinsics) is possible but it comes with performance degradations.

      For memory allocation, one call of malloc (and a corresponding call of free) would suffice. The temporary buffer has a size linear in the input integer and doesn't need to be a PyObject instance.

      Please let me know of other details that need to be worked out. Thanks!

    15. added
      3.13only security fixes
      and removed
      3.12only security fixes
      on Jan 5, 2024
    16. serhiy-storchaka commented on Jul 16, 2024

      @serhiy-storchaka
      Member

      Is there anything left to do in this issue or it can be closed?

    17. tim-one commented on Jul 16, 2024

      @tim-one
      MemberAuthor

      It's hard to tell because of the length and complexity of the discussion, but decimal.Decimal <-> int conversions remain quadratic-time. It's only str <-> int conversions that were effectively addressed.

    18. added and removed
      3.13only security fixes
      on Jul 12, 2025
    19. skirpichev commented on Sep 6, 2025

      @skirpichev
      Member

      decimal.Decimal <-> int conversion handled by libmpdec's mpd_qexport_/mpd_qimport_ API. We are going to remove libmpdec bundled copy in 3.16.

      Is CPython an appopriate place to fix/track libmpdec's issues? I think that fix should rather be submitted to upstream.

    20. tim-one commented on Sep 6, 2025

      @tim-one
      MemberAuthor

      If Python wants faster Decimal<->int conversion than libmpdec offers, then it's up to Python to arrange for that in its libmpdec wrapper. The "internal" in this issue's "Quadratic time internal base conversions" is there for a reason.

      Whether Stefan Krah wants to build faster conversion inside libmpdec is a different (although possibly related) issue. Making Python's story consistent across the Python types Python offers isn't really libmpdec's problem to solve. Sure, it would make life easier for us if libmpf\dec did do it much faster itself. But unless and until it does, this remains a Python issue.

    21. gpshead commented on Sep 6, 2025

      @gpshead
      Member

      closing as not planned, if someone is ever actually going to work on this they can reopen it with a plan.

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      extension-modulesC modules in the Modules dirinterpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagetype-featureA feature request or enhancement

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions