Skip to content

Provide C implementation for asyncio.current_task #100344

Description

@itamaro

Feature or enhancement

By providing a C implementation for asyncio.current_task, its performance can be improved.

Pitch

Performance improvement.

From Instagram profiling data, we've found that this function is called frequently, and a C implementation (in Cinder 3.8) showed more than 4x speedup in a microbenchmark.

Previous discussion

N/A

Linked PRs

Activity

  1. added a commit that references this issue on Dec 19, 2022
  2. kumaraditya303 commented on Dec 19, 2022

    @kumaraditya303
    Contributor

    Can you post benchmarks with pyperf?

  3. added
    performancePerformance or resource usage
    3.12only security fixes
    and removed
    type-featureA feature request or enhancement
    on Dec 19, 2022
  4. itamaro commented on Dec 19, 2022

    @itamaro
    ContributorAuthor

    Can you post benchmarks with pyperf?

    You mean pyperformance suite with and without the C acceleration?

  5. kumaraditya303 commented on Dec 19, 2022

    @kumaraditya303
    Contributor

    You mean pyperformance suite with and without the C acceleration?

    No this microbenchmark with pyperf.

  6. bluetech commented on Dec 19, 2022

    @bluetech
    Contributor

    As a note of support for making this faster, current_task is used in the asgiref.local implementation, which is used a lot by Django, and it shows up in profiles.

  7. itamaro commented on Dec 19, 2022

    @itamaro
    ContributorAuthor

    No this microbenchmark with pyperf.

    thanks for the clarification :)

    I don't know how to isolate the overhead of the event loop when using pyperf, so hope this is helpful:

    C implementation:

    $ python -m pyperf timeit -s 'from asyncio.tasks import _c_current_task as current_task' -s 'from asyncio import run' -s '
    async def main():
      for _ in range(10**6):
        current_task()
    ' 'run(main())'
    .....................
    Mean +- std dev: 33.4 ms +- 1.0 ms
    

    Python implementation:

    $ python -m pyperf timeit -s 'from asyncio.tasks import _py_current_task as current_task' -s 'from asyncio import run' -s '
    async def main():
      for _ in range(10**6):
        current_task()
    ' 'run(main())'
    .....................
    Mean +- std dev: 133 ms +- 8 ms
    
  8. kumaraditya303 commented on Dec 20, 2022

    @kumaraditya303
    Contributor

    The numbers looks interesting, it seems to be because there is no fastpath for dict.get like there is for list.append in ceval.

    @markshannon Do you have plans to optimize this?

  9. markshannon commented on Dec 20, 2022

    @markshannon
    Member

    @kumaraditya303 No plans at the moment.

    I'd be interested to see how this compared:

    def current_task(loop=None):
        """Return a currently executed task."""
        if loop is None:
            loop = events.get_running_loop()
        try:
            return _current_tasks[loop]
        except:
            return None

    I assume that current_task() is expected to return a task, not None, most of the time.

  10. itamaro commented on Dec 20, 2022

    @itamaro
    ContributorAuthor

    I assume that current_task() is expected to return a task, not None, most of the time.

    makes sense!

    this optimization make the python impl about 40% faster!

    $ python -m pyperf timeit -s 'from asyncio.tasks import _py_current_task as current_task' -s 'from asyncio import run' -s '
    async def main():
      for _ in range(10**6):
        current_task()
    ' 'run(main())'
    .....................
    Mean +- std dev: 81.1 ms +- 4.9 ms
    

    the C impl is still more than 2x faster than this, so maybe do both?

  11. gvanrossum commented on Dec 21, 2022

    @gvanrossum
    Member

    Why bother speeding up the Python version if we have the C version? There really aren't any interesting situations where the C accelerator is unavailable (that I know of).

  12. itamaro commented on Dec 21, 2022

    @itamaro
    ContributorAuthor

    Why bother speeding up the Python version if we have the C version? There really aren't any interesting situations where the C accelerator is unavailable (that I know of).

    In my mind it's "why not speed up the Python version?"
    I'm not aware of situations where it matters for cpython users, but maybe alternative implementations that use cpython's stdlib it would be valuable?
    anyway, I don't feel strongly about it. happy to revert to the existing python implementation if you'd prefer!

  13. gvanrossum commented on Dec 21, 2022

    @gvanrossum
    Member

    In my mind it's "why not speed up the Python version?"

    Because you're replacing one line of code with a well-known idiom with four lines of code that require the reader to follow carefully what's going on and why. For me, reading the version with .get() is much quicker than the try/except version.

    If we wrote hyper-optimized code like that everywhere, even when speed doesn't matter, we'd end up with considerably less readable code. So, in my mind the question very much needs to be "why speed it up".

  14. carljm commented on Dec 22, 2022

    @carljm
    Member

    The optimization in the Python version is also very much tuned for the current performance characteristics of the adaptive interpreter (i.e. that dict subscripts are much better optimized than dict method calls.) If alternate Python implementations are the only likely users of the Python implementation, this code change won't necessarily give similar speedups for them.

  15. added a commit that references this issue on Dec 22, 2022
  16. Repository owner moved this from Todo to Done in asyncioon Dec 22, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    • Status
      Done

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions