Repository navigation
Poor thread scaling when constructing instances or accessing attributes #139103
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Sep 18, 2025 - addedperformancePerformance or resource usagePerformance or resource usage
on Sep 18, 2025 - changed the title
[-]Poor thread scaling when constructing instances or accessings attributes[/-][+]Poor thread scaling when constructing instances or accessing attributes[/+]on Sep 18, 2025 I've started looking into this. The following is mostly for my own future reference, but may be useful if anyone else wants to take a look at this as well.
Here are the above examples added to
Tools/ftscalingbench/ftscalingbench.py: https://gist.github.com/colesbury/65234f8e8d9afad361981255d8259d7fDataclasses: there's reference count contention on the
__init__function. I think the function is created by callingexec()on a string, so the__init__function doesn't have deferred reference counting enabled. I think we can fix this by enabling deferred reference counting when the function is set on the type object intype_setattro.typing.NamedTuple: Calling
slot_tp_newdoesn't scale well.PyObject_GetAttr()increments the refcount and thePy_DECREFdecrements it. We want something like_PyObject_GetMethodStackRef()here that uses stackrefs so that we avoid reference count contention.Lines 10842 to 10856 in f0d8583
static PyObject * slot_tp_new(PyTypeObject *type, PyObject *args, PyObject *kwds) { PyThreadState *tstate = _PyThreadState_GET(); PyObject *func, *result; func = PyObject_GetAttr((PyObject *)type, &_Py_ID(__new__)); if (func == NULL) { return NULL; } result = _PyObject_Call_Prepend(tstate, func, (PyObject *)type, args, kwds); Py_DECREF(func); return result; } Enum: I think there's reference count contention on instances of
enum.property:
Line 173 in f0d8583
class property(DynamicClassAttribute): Hi @colesbury and @JukkaL ,
I'm interested in this topic. Could I have a try? 😊
Best Regards,
EdwardHi @colesbury ,
Sorry to bother you for this issue. 😊
I have read the dataclass code and find that it generated below code and runs with
exec.def __create_fn__(__dataclass_type_x__,__dataclass_HAS_DEFAULT_FACTORY__,__dataclass_builtins_object__,__dataclass___init___return_type__,__dataclasses_recursive_repr): def __init__(self,x:__dataclass_type_x__)->__dataclass___init___return_type__: self.x=x @__dataclasses_recursive_repr() def __repr__(self): return f"{self.__class__.__qualname__}(x={self.x!r})" def __eq__(self,other): if self is other: return True if other.__class__ is self.__class__: return self.x==other.x return NotImplemented return (__init__,__repr__,__eq__,)
Here is the reference code snippts.
Lines 492 to 524 in 2e5e6fd
# txt is the entire function we're going to execute, including the # bodies of the functions we're defining. Here's a greatly simplified # version: # def __create_fn__(): # def __init__(self, x, y): # self.x = x # self.y = y # @recursive_repr # def __repr__(self): # return f"cls(x={self.x!r},y={self.y!r})" # return __init__,__repr__ txt = f"def __create_fn__({local_vars}):\n{fns_src}\n return {return_names}" ns = {} exec(txt, self.globals, ns) fns = ns['__create_fn__'](**self.locals) # Now that we've generated the functions, assign them into cls. for name, fn in zip(self.names, fns): fn.__qualname__ = f"{cls.__qualname__}.{fn.__name__}" if self.unconditional_adds.get(name, False): setattr(cls, name, fn) else: already_exists = _set_new_attribute(cls, name, fn) # See if it's an error to overwrite this particular function. if already_exists and (msg_extra := self.overwrite_errors.get(name)): error_msg = (f'Cannot overwrite attribute {fn.__name__} ' f'in class {cls.__name__}') if not msg_extra is True: error_msg = f'{error_msg} {msg_extra}' raise TypeError(error_msg) So I create a case without dataclass below.
from threading import Thread from time import time def test_1(): class Foo: def __init__(self, x): self.x = x niter = 5 * 1000 * 1000 def benchmark(n): for i in range(n): Foo(x=1) for nth in (1, 4): t0 = time() threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)] for t in threads: t.start() for t in threads: t.join() print(f"{nth=} {(time() - t0) / nth}") def test_2(): class Foo2: def __init__(self, x): pass pass _Foo2_x = int create_str = """def create_init(_Foo2_x,): def __init__(self, x: _Foo2_x): self.x = x return (__init__,) """ ns = {} exec(create_str, globals(), ns) fn = ns['create_init']({**locals()}) setattr(Foo2, '__init__', fn[0]) niter = 5 * 1000 * 1000 def benchmark(n): for i in range(n): Foo2(x=1) for nth in (1, 4): t0 = time() threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)] for t in threads: t.start() for t in threads: t.join() print(f"{nth=} {(time() - t0) / nth}") if __name__ == "__main__": print("------test_1-------") test_1() print("------test_2-------") test_2()
But got below result.
------test_1------- nth=1 6.3677897453308105 nth=4 2.0216556191444397 ------test_2------- nth=1 6.5973052978515625 nth=4 2.26946008205413It seems that this case follow the dataclass logic to exec string and generate an
__init__.
But the performance didn't change.I use below configure to build cpython on main branch 1697cb5 .
./configure --with-pydebug --disable-gil "CC=clang"Could you help correct me on the understanding of dataclass's case? 😊
Wish you a good day!
Best Regards,
EdwardDon't use
--with-pydebugwhen benchmarking:uv run -p 3.14t python example.py:------test_1------- nth=1 1.1103649139404297 nth=4 0.2983349561691284 ------test_2------- nth=1 1.1990399360656738 nth=4 1.652941882610321Reacted by Edward XuDon't use
--with-pydebugwhen benchmarking:uv run -p 3.14t python example.py:------test_1------- nth=1 1.1103649139404297 nth=4 0.2983349561691284 ------test_2------- nth=1 1.1990399360656738 nth=4 1.652941882610321Thanks very much for your kind help and patience!
I will try to solve this issue. 😊- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Nov 15, 2025 typing.NamedTuple: Calling
slot_tp_newdoesn't scale well.PyObject_GetAttr()increments the refcount and thePy_DECREFdecrements it. We want something like_PyObject_GetMethodStackRef()here that uses stackrefs so that we avoid reference count contention.@colesbury In #141603 I used the
_PyObject_GetMethodStackRefapproach, but the scaling performance is unchanged. Not sure yet why that is.A small update: for
from collections import namedtuple Bar= namedtuple('Bar', ('x'))the call
B(1)is specialized to_CALL_NON_PY_GENERAL. Maybe that does not scale.Even with the modified
slot_tp_newthe scaling of getting an attribute is quite poor:Bar = namedtuple('Bar', ['x']) class Normal: pass @register_benchmark def attribute_normal_class(): for i in range(1000 * WORK_SCALE): Normal.__new__ @register_benchmark def attribute_namedtuple(): for i in range(1000 * WORK_SCALE): Bar.__new__attribute_normal_class 3.6x faster attribute_namedtuple 3.9x slower- added 3 commits that reference this issue
on Nov 19, 2025 7 remaining items
- added 4 commits that reference this issue
on Feb 3, 2026 - added 3 commits that reference this issue
on Feb 15, 2026 - added a commit that references this issue
on Feb 18, 2026 Also see this on various class/instance attributes in addition to the above PR, shall I be putting these in a new issue?
class MyClassWithInstanceAttr: def __init__(self): self.attr = object() # A single instance shared between threads. Reading its attribute (stored in # the inline __dict__ values, LOAD_ATTR_INSTANCE_VALUE) contends on the shared # value's reference count on every read. _shared_instance = MyClassWithInstanceAttr() @register_benchmark def shared_instance_attribute(): obj = _shared_instance for _ in range(1000 * WORK_SCALE): obj.attr obj.attr obj.attr./python_main.exe Tools/ftscalingbench/ftscalingbench.py shared_instance_attribute
Running benchmarks with 18 threads
shared_instance_attribute 25.0x slower
Bug report
Remaining scaling bugs
Bug description:
When constructing dataclass or NamedTuple instances on multiple threads (on a free threading build), or accessing enum class attributes, performance doesn't scale when using multiple threads.
Regular class example (scales well):
Dataclass example (doesn't scale well):
Named tuple example (doesn't scale well):
Enum example (doesn't scale well):
Results on recent main branch (running on an EC2 instance):
The expected behavior is that when using 4 threads (
nth=4), the elapsed time per benchmark iteration (the second printed value) goes down significantly compared to when using a single thread (nth=1), which happens with the first benchmark (b_regular_class.py) but not the others.cc @colesbury (we discussed this at CPython Core Dev Sprint in person)
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs
dataclass.__init__perf issue #141596dataclass.__init__perf issue (gh-141596) #141750