Skip to content

_colorize module is slow to import, affecting traceback and logging modules #144384

Description

@danielhollas

“It is a truth universally acknowledged, that a module in possession of great colors, suffers import time slowdown.

_colorize module is slow to import due to its use of dataclasses

python -Ximporttime -c "import _colorize"

Visualization in tuna

Image

Notice that significant time is also spent executing the _colorize, due to the creation of its dataclasses

❯ uvx python@3.15 -m cProfile -m _colorize | head -15
         5609 function calls (5489 primitive calls) in 0.006 seconds

   Ordered by: cumulative time

   ncalls  tottime  percall  cumtime  percall filename:lineno(function)
      9/1    0.003    0.000    0.006    0.006 {built-in method builtins.exec}
        1    0.000    0.000    0.006    0.006 <string>:1(<module>)
        1    0.000    0.000    0.006    0.006 <frozen runpy>:199(run_module)
        1    0.000    0.000    0.006    0.006 <frozen runpy>:65(_run_code)
        1    0.000    0.000    0.006    0.006 _colorize.py:1(<module>)
        7    0.000    0.000    0.005    0.001 dataclasses.py:1432(wrap)
        7    0.000    0.000    0.005    0.001 dataclasses.py:986(_process_class)
        7    0.000    0.000    0.003    0.000 dataclasses.py:478(add_fns_to_class)
        4    0.000    0.000    0.001    0.000 inspect.py:3342(signature)
        4    0.000    0.000    0.001    0.000 inspect.py:3055(from_callable)

The slow import time directly affects (among others) the traceback module which in turn affects logging, which is 40% slower in 3.14 and 3.15 compared to 3.13. That is quite unfortunate since in applications that care about import time, logging module is typically hard to avoid (I originally discovered this when looking at pip's startup time)

 hyperfine -w 10 "python3.13 -c 'import logging'"  "python3.15 -c 'import logging'" 
Benchmark 1: python3.13 -c 'import logging'
  Time (mean ± σ):      21.0 ms ±   3.4 ms    [User: 14.8 ms, System: 5.9 ms]
  Range (min … max):    16.2 ms …  28.2 ms    155 runs
 
Benchmark 2: python3.15 -c 'import logging'
  Time (mean ± σ):      29.6 ms ±   2.1 ms    [User: 22.2 ms, System: 7.0 ms]
  Range (min … max):    26.4 ms …  35.1 ms    100 runs
 
Summary
  python3.13 -c 'import logging' ran
    1.40 ± 0.25 times faster than /home/hollas/.local/share/uv/python/cpython-3.15.0a5-linux-x86_64-gnu/bin/python3.15 -c 'import logging'

It's not very clear how to make this better (besides not using dataclasses). We tried to make traceback lazy in logging, but failed. #112995

Linked PRs

Activity

  1. added
    performancePerformance or resource usage
    stdlibStandard Library Python modules in the Lib/ directory
    3.14bugs and security fixes
    3.15pre-release feature fixes, bugs and security fixes
    on Feb 2, 2026
  2. picnixz commented on Feb 2, 2026

    @picnixz
    Member
  3. picnixz commented on Feb 2, 2026

    @picnixz
    Member

    So the problem is colorize being imported by traceback which is imported by logging right? why not making colorize lazy in traceback?

  4. picnixz commented on Feb 2, 2026

    @picnixz
    Member

    OTOH, what if instead of a dataclass we use a namedtuple? (not from typing but from collections) or a custom class with slots? While we would lose kw-only support with namedtuples, we can backport the change and avoid changing other modules.

  5. added
    type-bugAn unexpected behavior, bug, or error
    and removed
    type-featureA feature request or enhancement
    on Feb 2, 2026
  6. picnixz commented on Feb 2, 2026

    @picnixz
    Member

    @StanFromIreland The fact that logging is affected makes it a bug for me as logging may be quite frequent in applications and those applications may also want to retain small startup time.

  7. danielhollas commented on Feb 2, 2026

    @danielhollas
    ContributorAuthor

    why not making colorize lazy in traceback?

    I looked at that and it didn't seem easy as it is used in a bunch of places. But once the lazy import PR lands, it seems doable. But I am not sure if lazy importing in the traceback module is a great idea in general (e.g. you probably don't want to be importing modules in the middle of handling exceptions?).

  8. picnixz commented on Feb 2, 2026

    @picnixz
    Member

    Unless I am mistaken, traceback is only meant for rendering tracebacks, so it should be fine. Lazy imports will not be available for 3.14 so it may be a good solution for the backport. How slower is the logging import?

  9. ZeroIntensity commented on Feb 2, 2026

    @ZeroIntensity
    Member

    How difficult would it be to speed up dataclass creation in general? I have two ideas:

    1. Add a C accelerator for dataclasses. This might be maintenance-heavy.
    2. Add a way to lazily construct a dataclass, so parsing of the class and whatnot isn't done until __new__ is called for the first time.

    Dataclasses are becoming increasingly common, and I think it would be nice if they weren't such a significant performance hit at import time.

  10. picnixz commented on Feb 2, 2026

    @picnixz
    Member

    I do not think the problem is @dataclass; the problem is that we import a lot of modules in the dataclasses module:

    import re
    import sys
    import copy
    import types
    import inspect
    import keyword
    import itertools
    import annotationlib
    import abc

    We import inspect and enum and those are heavy modules.

  11. 5 remaining items

  12. hugovk commented on Feb 16, 2026

    @hugovk
    Member

    Hugo said he would check this assumption against the most popular packages on PyPI.

    Searching the top 15k PyPI packages from June (actually 14,041 with downloadable source), 2,485 match the "(import dataclasses)|(from dataclasses import)" regex, or 17.7%.

  13. hugovk commented on Feb 16, 2026

    @hugovk
    Member

    See also #144387 to lazy import inspect in dataclasses.

  14. DavidCEllis commented on Feb 17, 2026

    @DavidCEllis
    Contributor

    How difficult would it be to speed up dataclass creation in general? I have two ideas:

    1. Add a C accelerator for `dataclasses`. This might be maintenance-heavy.
    
    2. Add a way to lazily construct a dataclass, so parsing of the class and whatnot isn't done until `__new__` is called for the first time.
    

    Dataclasses are becoming increasingly common, and I think it would be nice if they weren't such a significant performance hit at import time.

    Outside of the import time improvements, most of the construction time in dataclasses is spent on the exec call to create all of the functions.

    from cProfile import Profile
    from pstats import Stats
    from dataclasses import dataclass
    
    with Profile() as p:
        for _ in range(5000):
            @dataclass
            class Example:
                a: int
                b: str
                c: str
                d: list[int]
                e: dict[str, int]
    
        stats = Stats(p)
        stats.strip_dirs()
        stats.sort_stats("tottime")
        stats.print_stats(10)

    Note: This is run on top of #144387 - without that you'll also see inspect in here for docstring creation

       ncalls  tottime  percall  cumtime  percall filename:lineno(function)
         5000    0.421    0.000    0.422    0.000 {built-in method builtins.exec}
         5000    0.094    0.000    0.750    0.000 dataclasses.py:1014(_process_class)
        25000    0.032    0.000    0.058    0.000 dataclasses.py:831(_get_field)
         5000    0.027    0.000    0.480    0.000 dataclasses.py:477(add_fns_to_class)
         5000    0.021    0.000    0.023    0.000 {built-in method builtins.__build_class__}
         5000    0.018    0.000    0.043    0.000 dataclasses.py:669(_init_fn)
        15000    0.015    0.000    0.023    0.000 dataclasses.py:445(add_fn)
       180001    0.015    0.000    0.015    0.000 {built-in method builtins.isinstance}
        90000    0.009    0.000    0.009    0.000 {built-in method builtins.getattr}
        55000    0.009    0.000    0.009    0.000 {method 'join' of 'str' objects}
    
    

    One way to potentially make this faster in many cases would be to defer the generation of the methods using descriptors, only generating the methods that are actually used. It would be slightly slower for classes that use all methods but faster if some are never used1. This is how I generate the methods in my own classbuilder.

    Footnotes

    1. I suspect a lot of dataclasses never actually use __repr__ at runtime. ↩

  15. FFY00 commented on Apr 1, 2026

    @FFY00
    Member

    I agree with @ambv about improving dataclasses instead, though I'd like to point out that an easy solution to speed up the type construction would be to freeze the _colorize module.

  16. hugovk commented on Apr 2, 2026

    @hugovk
    Member

    Freezing _colorize doesn't make much difference:

    Before After
    Image Image

    About two thirds of the time is spent importing dataclasses, and about a third is creating the dataclass for each theme.

  17. DavidCEllis commented on Apr 10, 2026

    @DavidCEllis
    Contributor

    I'd like to add that while this issue mentions traceback and logging, _colorize also makes argparse signifcantly slower. It's not imported at top level, but it is imported as soon as you actually create a parser (color=False does not change this).

    Comparing 3.13.13 and 3.15.0a8:

    Benchmark 1: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      35.1 ms ±   4.9 ms    [User: 27.9 ms, System: 7.0 ms]
      Range (min … max):    26.0 ms …  44.9 ms    30 runs
     
    Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      61.9 ms ±   3.8 ms    [User: 52.2 ms, System: 9.3 ms]
      Range (min … max):    53.9 ms …  73.5 ms    30 runs
     
    Summary
      .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran
        1.76 ± 0.27 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
    

    Not all of this is colorize (import shutil also got a little slower due to the addition of zstd), but most of it is. If you don't create a parser, the import time looks the same but in practice it's gotten a lot slower.

  18. danielhollas commented on Apr 10, 2026

    @danielhollas
    ContributorAuthor

    @DavidCEllis interesting. Can you try measuring how much #144387 helps here?

  19. DavidCEllis commented on Apr 10, 2026

    @DavidCEllis
    Contributor

    lazy_dataclasses is dataclasses from that PR, with an extra line to replace sys.modules["dataclasses"] so the time should be correct for that PR.

    Benchmark 1: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      35.5 ms ±   4.4 ms    [User: 28.4 ms, System: 7.0 ms]
      Range (min … max):    26.6 ms …  43.9 ms    30 runs
     
    Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      61.9 ms ±   3.9 ms    [User: 53.3 ms, System: 8.5 ms]
      Range (min … max):    52.7 ms …  68.2 ms    30 runs
     
    Benchmark 3: .venv_315/bin/python -c 'import lazy_dataclasses; import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      52.3 ms ±   2.9 ms    [User: 44.0 ms, System: 8.2 ms]
      Range (min … max):    46.5 ms …  58.9 ms    30 runs
     
    Summary
      .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran
        1.47 ± 0.20 times faster than .venv_315/bin/python -c 'import lazy_dataclasses; import argparse; argparse.ArgumentParser()'
        1.74 ± 0.24 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
    

    It's definitely better, but dataclasses is still getting hit by annotationlib which then hits argparse.

  20. hugovk commented on Apr 11, 2026

    @hugovk
    Member

    @DavidCEllis What OS is that with?

    On macOS, 3.13 and 3.15 are much closer. Here is using the official installers (3.13.13, 3.15.0a8) and a local build of main with optimisations:

    ❯ hyperfine \
     "python3.13  -c 'import argparse; argparse.ArgumentParser()'" \
     "python3.15  -c 'import argparse; argparse.ArgumentParser()'" \
    "./python.exe -c 'import argparse; argparse.ArgumentParser()'"
    Benchmark 1: python3.13  -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      23.7 ms ±   1.2 ms    [User: 17.7 ms, System: 5.0 ms]
      Range (min … max):    22.2 ms …  28.3 ms    112 runs
    
    Benchmark 2: python3.15  -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      32.2 ms ±   0.8 ms    [User: 25.9 ms, System: 5.2 ms]
      Range (min … max):    31.0 ms …  35.4 ms    81 runs
    
    Benchmark 3: ./python.exe -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      29.4 ms ±   1.2 ms    [User: 24.3 ms, System: 4.2 ms]
      Range (min … max):    28.2 ms …  38.6 ms    71 runs
    
      Warning: The first benchmarking run for this command was significantly slower than the rest (38.6 ms). This could be caused by (filesystem) caches that were not filled until after the first run. You should consider using the '--warmup' option to fill those caches before the actual benchmark. Alternatively, use the '--prepare' option to clear the caches before each timing run.
    
    Summary
      python3.13  -c 'import argparse; argparse.ArgumentParser()' ran
        1.24 ± 0.08 times faster than ./python.exe -c 'import argparse; argparse.ArgumentParser()'
        1.36 ± 0.07 times faster than python3.15  -c 'import argparse; argparse.ArgumentParser()'

    And #144387 is a big improvement, bringing it even closer to 3.13:

    ❯ hyperfine \
     "python3.13  -c 'import argparse; argparse.ArgumentParser()'" \
     "python3.15  -c 'import argparse; argparse.ArgumentParser()'" \
    "./python.exe -c 'import argparse; argparse.ArgumentParser()'"
    Benchmark 1: python3.13  -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      23.1 ms ±   0.5 ms    [User: 17.4 ms, System: 4.8 ms]
      Range (min … max):    22.3 ms …  25.2 ms    105 runs
    
    Benchmark 2: python3.15  -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      32.1 ms ±   0.8 ms    [User: 25.9 ms, System: 5.2 ms]
      Range (min … max):    30.9 ms …  35.8 ms    80 runs
    
    Benchmark 3: ./python.exe -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      24.2 ms ±   0.5 ms    [User: 19.7 ms, System: 3.7 ms]
      Range (min … max):    23.2 ms …  27.7 ms    96 runs
    
    Summary
      python3.13  -c 'import argparse; argparse.ArgumentParser()' ran
        1.05 ± 0.03 times faster than ./python.exe -c 'import argparse; argparse.ArgumentParser()'
        1.39 ± 0.05 times faster than python3.15  -c 'import argparse; argparse.ArgumentParser()'
  21. DavidCEllis commented on Apr 11, 2026

    @DavidCEllis
    Contributor

    It's a not exactly new laptop running Ubuntu 24.04, using the builds from uv. Times on this machine are slower, but generally more consistent than on the faster desktop I have running Fedora.

    I'm actually just noting that on looking more closely - with the patch - it's less of an impact of the ast import and more _colorize itself being surprisingly slow? Is that all just exec calls for dataclasses? I'll check on the faster machine early next week to see if my numbers are more like yours.

  22. hugovk commented on Apr 11, 2026

    @hugovk
    Member

    I'm actually just noting that on looking more closely - with the patch - it's less of an impact of the ast import and more _colorize itself being surprisingly slow? Is that all just exec calls for dataclasses?

    Yes, I think so. See #144384 (comment) -- 2/3 importing dataclasses, 1/3 execution.

    And see #144387 (comment) to show improvements for _colorize with the PR over main.

  23. danielhollas commented on Apr 11, 2026

    @danielhollas
    ContributorAuthor

    Is that all just exec calls for dataclasses?

    Yes, that can be seen if your run the import under cProfile.

  24. DavidCEllis commented on Apr 13, 2026

    @DavidCEllis
    Contributor

    I see better numbers on a faster machine, but still not as close to the 3.13 numbers as @hugovk

    Benchmark 1: .venv_lazyinspect/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      21.3 ms ±   3.7 ms    [User: 18.3 ms, System: 2.9 ms]
      Range (min … max):    18.5 ms …  30.7 ms    50 runs
     
    Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      26.1 ms ±   3.3 ms    [User: 21.9 ms, System: 4.1 ms]
      Range (min … max):    21.1 ms …  32.7 ms    50 runs
     
    Benchmark 3: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()'
      Time (mean ± σ):      17.4 ms ±   3.4 ms    [User: 14.2 ms, System: 3.1 ms]
      Range (min … max):    11.9 ms …  29.1 ms    50 runs
     
    Summary
      .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran
        1.23 ± 0.32 times faster than .venv_lazyinspect/bin/python -c 'import argparse; argparse.ArgumentParser()'
        1.50 ± 0.35 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'
    

    Is it worth making a separate issue for argparse to defer _colorize further (until it actually needs to colour something)?

  25. hugovk commented on May 6, 2026

    @hugovk
    Member

    I opened #149318 to lazily import _colorize, which saves around 8ms and makes importing some modules between 17% and 80% faster.

  26. added a commit that references this issue on May 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.14bugs and security fixes3.15pre-release feature fixes, bugs and security fixesperformancePerformance or resource usagestdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions