Repository navigation
Provide C implementation for asyncio.current_task #100344
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Dec 19, 2022 Can you post benchmarks with pyperf?
- addedperformancePerformance or resource usagePerformance or resource usage3.12only security fixesonly security fixesand removedtype-featureA feature request or enhancementA feature request or enhancement
on Dec 19, 2022 Can you post benchmarks with pyperf?
You mean pyperformance suite with and without the C acceleration?
You mean pyperformance suite with and without the C acceleration?
No this microbenchmark with pyperf.
As a note of support for making this faster,
current_taskis used in theasgiref.localimplementation, which is used a lot by Django, and it shows up in profiles.Reacted by Itamar OrenNo this microbenchmark with pyperf.
thanks for the clarification :)
I don't know how to isolate the overhead of the event loop when using pyperf, so hope this is helpful:
C implementation:
$ python -m pyperf timeit -s 'from asyncio.tasks import _c_current_task as current_task' -s 'from asyncio import run' -s ' async def main(): for _ in range(10**6): current_task() ' 'run(main())' ..................... Mean +- std dev: 33.4 ms +- 1.0 msPython implementation:
$ python -m pyperf timeit -s 'from asyncio.tasks import _py_current_task as current_task' -s 'from asyncio import run' -s ' async def main(): for _ in range(10**6): current_task() ' 'run(main())' ..................... Mean +- std dev: 133 ms +- 8 msReacted by Guido van RossumThe numbers looks interesting, it seems to be because there is no fastpath for
dict.getlike there is forlist.appendin ceval.@markshannon Do you have plans to optimize this?
Reacted by Itamar Oren@kumaraditya303 No plans at the moment.
I'd be interested to see how this compared:
def current_task(loop=None): """Return a currently executed task.""" if loop is None: loop = events.get_running_loop() try: return _current_tasks[loop] except: return None
I assume that
current_task()is expected to return a task, notNone, most of the time.Reacted by Itamar OrenI assume that
current_task()is expected to return a task, notNone, most of the time.makes sense!
this optimization make the python impl about 40% faster!
$ python -m pyperf timeit -s 'from asyncio.tasks import _py_current_task as current_task' -s 'from asyncio import run' -s ' async def main(): for _ in range(10**6): current_task() ' 'run(main())' ..................... Mean +- std dev: 81.1 ms +- 4.9 msthe C impl is still more than 2x faster than this, so maybe do both?
Why bother speeding up the Python version if we have the C version? There really aren't any interesting situations where the C accelerator is unavailable (that I know of).
Why bother speeding up the Python version if we have the C version? There really aren't any interesting situations where the C accelerator is unavailable (that I know of).
In my mind it's "why not speed up the Python version?"
I'm not aware of situations where it matters for cpython users, but maybe alternative implementations that use cpython's stdlib it would be valuable?
anyway, I don't feel strongly about it. happy to revert to the existing python implementation if you'd prefer!In my mind it's "why not speed up the Python version?"
Because you're replacing one line of code with a well-known idiom with four lines of code that require the reader to follow carefully what's going on and why. For me, reading the version with
.get()is much quicker than thetry/exceptversion.If we wrote hyper-optimized code like that everywhere, even when speed doesn't matter, we'd end up with considerably less readable code. So, in my mind the question very much needs to be "why speed it up".
Reacted by Itamar Oren, Paweł Rubin and AlenPaulVargheseThe optimization in the Python version is also very much tuned for the current performance characteristics of the adaptive interpreter (i.e. that dict subscripts are much better optimized than dict method calls.) If alternate Python implementations are the only likely users of the Python implementation, this code change won't necessarily give similar speedups for them.
Reacted by Itamar Oren and Asif Saif Uddin {"Auvi":"অভি"}- added a commit that references this issue
on Dec 27, 2022 - added a commit that references this issue
on Dec 28, 2022
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Feature or enhancement
By providing a C implementation for
asyncio.current_task, its performance can be improved.Pitch
Performance improvement.
From Instagram profiling data, we've found that this function is called frequently, and a C implementation (in Cinder 3.8) showed more than 4x speedup in a microbenchmark.
Previous discussion
N/A
Linked PRs