Repository navigation
_colorize module is slow to import, affecting traceback and logging modules #144384
Description
Activity
- addedperformancePerformance or resource usagePerformance or resource usagestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directory3.14bugs and security fixesbugs and security fixes3.15pre-release feature fixes, bugs and security fixespre-release feature fixes, bugs and security fixes
on Feb 2, 2026 cc @hugovk
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Feb 2, 2026 So the problem is colorize being imported by traceback which is imported by logging right? why not making colorize lazy in traceback?
OTOH, what if instead of a dataclass we use a namedtuple? (not from typing but from collections) or a custom class with slots? While we would lose kw-only support with namedtuples, we can backport the change and avoid changing other modules.
Reacted by Alex Waygood- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errorand removedtype-featureA feature request or enhancementA feature request or enhancement
on Feb 2, 2026 @StanFromIreland The fact that logging is affected makes it a bug for me as logging may be quite frequent in applications and those applications may also want to retain small startup time.
Reacted by Stan Ulbrychwhy not making colorize lazy in traceback?
I looked at that and it didn't seem easy as it is used in a bunch of places. But once the lazy import PR lands, it seems doable. But I am not sure if lazy importing in the
tracebackmodule is a great idea in general (e.g. you probably don't want to be importing modules in the middle of handling exceptions?).Unless I am mistaken, traceback is only meant for rendering tracebacks, so it should be fine. Lazy imports will not be available for 3.14 so it may be a good solution for the backport. How slower is the logging import?
How difficult would it be to speed up dataclass creation in general? I have two ideas:
- Add a C accelerator for
dataclasses. This might be maintenance-heavy. - Add a way to lazily construct a dataclass, so parsing of the class and whatnot isn't done until
__new__is called for the first time.
Dataclasses are becoming increasingly common, and I think it would be nice if they weren't such a significant performance hit at import time.
Reacted by Łukasz Langa and M. Kocher- Add a C accelerator for
I do not think the problem is
@dataclass; the problem is that we import a lot of modules in the dataclasses module:import re import sys import copy import types import inspect import keyword import itertools import annotationlib import abc
We import
inspectandenumand those are heavy modules.5 remaining items
Hugo said he would check this assumption against the most popular packages on PyPI.
Searching the top 15k PyPI packages from June (actually 14,041 with downloadable source), 2,485 match the
"(import dataclasses)|(from dataclasses import)"regex, or 17.7%.See also #144387 to lazy import
inspectindataclasses.How difficult would it be to speed up dataclass creation in general? I have two ideas:
1. Add a C accelerator for `dataclasses`. This might be maintenance-heavy. 2. Add a way to lazily construct a dataclass, so parsing of the class and whatnot isn't done until `__new__` is called for the first time.Dataclasses are becoming increasingly common, and I think it would be nice if they weren't such a significant performance hit at import time.
Outside of the import time improvements, most of the construction time in
dataclassesis spent on theexeccall to create all of the functions.from cProfile import Profile from pstats import Stats from dataclasses import dataclass with Profile() as p: for _ in range(5000): @dataclass class Example: a: int b: str c: str d: list[int] e: dict[str, int] stats = Stats(p) stats.strip_dirs() stats.sort_stats("tottime") stats.print_stats(10)
Note: This is run on top of #144387 - without that you'll also see
inspectin here for docstring creationncalls tottime percall cumtime percall filename:lineno(function) 5000 0.421 0.000 0.422 0.000 {built-in method builtins.exec} 5000 0.094 0.000 0.750 0.000 dataclasses.py:1014(_process_class) 25000 0.032 0.000 0.058 0.000 dataclasses.py:831(_get_field) 5000 0.027 0.000 0.480 0.000 dataclasses.py:477(add_fns_to_class) 5000 0.021 0.000 0.023 0.000 {built-in method builtins.__build_class__} 5000 0.018 0.000 0.043 0.000 dataclasses.py:669(_init_fn) 15000 0.015 0.000 0.023 0.000 dataclasses.py:445(add_fn) 180001 0.015 0.000 0.015 0.000 {built-in method builtins.isinstance} 90000 0.009 0.000 0.009 0.000 {built-in method builtins.getattr} 55000 0.009 0.000 0.009 0.000 {method 'join' of 'str' objects}One way to potentially make this faster in many cases would be to defer the generation of the methods using descriptors, only generating the methods that are actually used. It would be slightly slower for classes that use all methods but faster if some are never used1. This is how I generate the methods in my own classbuilder.
Footnotes
-
I suspect a lot of dataclasses never actually use
__repr__at runtime. ↩
-
I agree with @ambv about improving
dataclassesinstead, though I'd like to point out that an easy solution to speed up the type construction would be to freeze the_colorizemodule.I'd like to add that while this issue mentions
tracebackandlogging,_colorizealso makesargparsesignifcantly slower. It's not imported at top level, but it is imported as soon as you actually create a parser (color=Falsedoes not change this).Comparing 3.13.13 and 3.15.0a8:
Benchmark 1: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 35.1 ms ± 4.9 ms [User: 27.9 ms, System: 7.0 ms] Range (min … max): 26.0 ms … 44.9 ms 30 runs Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 61.9 ms ± 3.8 ms [User: 52.2 ms, System: 9.3 ms] Range (min … max): 53.9 ms … 73.5 ms 30 runs Summary .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran 1.76 ± 0.27 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'Not all of this is colorize (
import shutilalso got a little slower due to the addition of zstd), but most of it is. If you don't create a parser, the import time looks the same but in practice it's gotten a lot slower.@DavidCEllis interesting. Can you try measuring how much #144387 helps here?
lazy_dataclassesis dataclasses from that PR, with an extra line to replacesys.modules["dataclasses"]so the time should be correct for that PR.Benchmark 1: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 35.5 ms ± 4.4 ms [User: 28.4 ms, System: 7.0 ms] Range (min … max): 26.6 ms … 43.9 ms 30 runs Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 61.9 ms ± 3.9 ms [User: 53.3 ms, System: 8.5 ms] Range (min … max): 52.7 ms … 68.2 ms 30 runs Benchmark 3: .venv_315/bin/python -c 'import lazy_dataclasses; import argparse; argparse.ArgumentParser()' Time (mean ± σ): 52.3 ms ± 2.9 ms [User: 44.0 ms, System: 8.2 ms] Range (min … max): 46.5 ms … 58.9 ms 30 runs Summary .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran 1.47 ± 0.20 times faster than .venv_315/bin/python -c 'import lazy_dataclasses; import argparse; argparse.ArgumentParser()' 1.74 ± 0.24 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'It's definitely better, but
dataclassesis still getting hit byannotationlibwhich then hitsargparse.@DavidCEllis What OS is that with?
On macOS, 3.13 and 3.15 are much closer. Here is using the official installers (3.13.13, 3.15.0a8) and a local build of
mainwith optimisations:❯ hyperfine \ "python3.13 -c 'import argparse; argparse.ArgumentParser()'" \ "python3.15 -c 'import argparse; argparse.ArgumentParser()'" \ "./python.exe -c 'import argparse; argparse.ArgumentParser()'" Benchmark 1: python3.13 -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 23.7 ms ± 1.2 ms [User: 17.7 ms, System: 5.0 ms] Range (min … max): 22.2 ms … 28.3 ms 112 runs Benchmark 2: python3.15 -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 32.2 ms ± 0.8 ms [User: 25.9 ms, System: 5.2 ms] Range (min … max): 31.0 ms … 35.4 ms 81 runs Benchmark 3: ./python.exe -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 29.4 ms ± 1.2 ms [User: 24.3 ms, System: 4.2 ms] Range (min … max): 28.2 ms … 38.6 ms 71 runs Warning: The first benchmarking run for this command was significantly slower than the rest (38.6 ms). This could be caused by (filesystem) caches that were not filled until after the first run. You should consider using the '--warmup' option to fill those caches before the actual benchmark. Alternatively, use the '--prepare' option to clear the caches before each timing run. Summary python3.13 -c 'import argparse; argparse.ArgumentParser()' ran 1.24 ± 0.08 times faster than ./python.exe -c 'import argparse; argparse.ArgumentParser()' 1.36 ± 0.07 times faster than python3.15 -c 'import argparse; argparse.ArgumentParser()'
And #144387 is a big improvement, bringing it even closer to 3.13:
❯ hyperfine \ "python3.13 -c 'import argparse; argparse.ArgumentParser()'" \ "python3.15 -c 'import argparse; argparse.ArgumentParser()'" \ "./python.exe -c 'import argparse; argparse.ArgumentParser()'" Benchmark 1: python3.13 -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 23.1 ms ± 0.5 ms [User: 17.4 ms, System: 4.8 ms] Range (min … max): 22.3 ms … 25.2 ms 105 runs Benchmark 2: python3.15 -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 32.1 ms ± 0.8 ms [User: 25.9 ms, System: 5.2 ms] Range (min … max): 30.9 ms … 35.8 ms 80 runs Benchmark 3: ./python.exe -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 24.2 ms ± 0.5 ms [User: 19.7 ms, System: 3.7 ms] Range (min … max): 23.2 ms … 27.7 ms 96 runs Summary python3.13 -c 'import argparse; argparse.ArgumentParser()' ran 1.05 ± 0.03 times faster than ./python.exe -c 'import argparse; argparse.ArgumentParser()' 1.39 ± 0.05 times faster than python3.15 -c 'import argparse; argparse.ArgumentParser()'
It's a not exactly new laptop running Ubuntu 24.04, using the builds from
uv. Times on this machine are slower, but generally more consistent than on the faster desktop I have running Fedora.I'm actually just noting that on looking more closely - with the patch - it's less of an impact of the
astimport and more_colorizeitself being surprisingly slow? Is that all justexeccalls for dataclasses? I'll check on the faster machine early next week to see if my numbers are more like yours.I'm actually just noting that on looking more closely - with the patch - it's less of an impact of the
astimport and more_colorizeitself being surprisingly slow? Is that all justexeccalls for dataclasses?Yes, I think so. See #144384 (comment) -- 2/3 importing
dataclasses, 1/3 execution.And see #144387 (comment) to show improvements for
_colorizewith the PR overmain.Is that all just exec calls for dataclasses?
Yes, that can be seen if your run the import under cProfile.
I see better numbers on a faster machine, but still not as close to the 3.13 numbers as @hugovk
- .venv_lazyinspect - build from gh-137855: Lazy import
inspectmodule in dataclasses #144387 with optimisations. - .venv_315 - Python 3.15.0a8 from uv
- .venv_313 - Python 3.13.13 from uv
Benchmark 1: .venv_lazyinspect/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 21.3 ms ± 3.7 ms [User: 18.3 ms, System: 2.9 ms] Range (min … max): 18.5 ms … 30.7 ms 50 runs Benchmark 2: .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 26.1 ms ± 3.3 ms [User: 21.9 ms, System: 4.1 ms] Range (min … max): 21.1 ms … 32.7 ms 50 runs Benchmark 3: .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' Time (mean ± σ): 17.4 ms ± 3.4 ms [User: 14.2 ms, System: 3.1 ms] Range (min … max): 11.9 ms … 29.1 ms 50 runs Summary .venv_313/bin/python -c 'import argparse; argparse.ArgumentParser()' ran 1.23 ± 0.32 times faster than .venv_lazyinspect/bin/python -c 'import argparse; argparse.ArgumentParser()' 1.50 ± 0.35 times faster than .venv_315/bin/python -c 'import argparse; argparse.ArgumentParser()'Is it worth making a separate issue for
argparseto defer_colorizefurther (until it actually needs to colour something)?- .venv_lazyinspect - build from gh-137855: Lazy import
I opened #149318 to lazily import
_colorize, which saves around 8ms and makes importing some modules between 17% and 80% faster.Reacted by Daniel Hollas- added a commit that references this issue
on May 6, 2026


_colorizemodule is slow to import due to its use of dataclassespython -Ximporttime -c "import _colorize"Visualization in tuna
Notice that significant time is also spent executing the
_colorize, due to the creation of its dataclassesThe slow import time directly affects (among others) the
tracebackmodule which in turn affectslogging, which is 40% slower in 3.14 and 3.15 compared to 3.13. That is quite unfortunate since in applications that care about import time,loggingmodule is typically hard to avoid (I originally discovered this when looking at pip's startup time)It's not very clear how to make this better (besides not using dataclasses). We tried to make traceback lazy in logging, but failed. #112995
Linked PRs
_colorizeperformance by replacingdataclasswith regular class #144879_colorize#149318