Repository navigation
Performance regression 3.10b1: inlining issue in the big _PyEval_EvalFrameDefault() function with Visual Studio (MSC) #89279
Description
Activity
pyperformance on Windows shows some gap between 3.10a7 and 3.10b1.
The following are the ratios compared with 3.10a7 (the higher the slower).-------------------------------------------------
Windows x64 | PGO release official-binary
----------------+--------------------------------
20210405 |
3.10a7 | 1.00 1.24 1.00 (PGO?)
20210408-07:58 |
b98eba5 | 0.98
20210408-10:22 |- PR25244 | 1.04
20210503 |
3.10b1 | 1.07 1.21 1.07
-------------------------------------------------
Windows x86 | PGO release official-binary
----------------+--------------------------------
20210405 |
3.10a7 | 1.00 1.25 1.27 (release?)
20210408-07:58 |
b98eba5 | 1.00
20210408-10:22 | - PR25244 | 1.11
20210503 |
3.10b1 | 1.14 1.28 1.29
Since PR25244 (28d28e0),
_PyEval_EvalFrameDefault() in ceval.c has seemed to be unoptimized with PGO (msvc14.29.16.10).
At least the functions below have become un-inlined there at all.(1) _Py_DECREF() (from Py_DECREF,Py_CLEAR,Py_SETREF)
(2) _Py_XDECREF() (from Py_XDECREF,SETLOCAL)
(3) _Py_IS_TYPE() (from PyXXX_CheckExact)
(4) _Py_atomic_load_32bit_impl() (from CHECK_EVAL_BREAKER)I tried in vain other linker options like thread-safe-profiling, agressive-code-generation, /OPT:NOREF.
3.10a7 can inline them in the eval-loop even if profiling only test_array.py.I measured overheads of (1)~(4) on my own build whose eval-loop uses macros instead of them.
-----------------------------------------------------------------
Windows x64 | PGO patched overhead in eval-loop
----------------+------------------------------------------------
3.10a7 | 1.00
20210802 |
3.10rc1 | 1.09 1.05 4% (slow 43, fast 5, same 10)
20210831-20:42 |
863154c | 0.95 0.90 5% (slow 48, fast 3, same 7)
(3.11a0+) |
-----------------------------------------------------------------
Windows x86 | PGO patched overhead in eval-loop
----------------+------------------------------------------------
3.10a7 | 1.00
20210802 |
3.10rc1 | 1.15 1.13 2% (slow 29, fast 14, same 15)
20210831-20:42 |
863154c | 1.05 1.02 3% (slow 44, fast 7, same 7)
(3.11a0+) |- PR25244 | 1.04
- addedperformancePerformance or resource usagePerformance or resource usage3.10 (EOL)end of lifeend of life3.11only security fixesonly security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Sep 6, 2021 Rather than defining again functions as macro, you should consider using __forceinline function attribute: see bpo-45094.
Perhaps these critical code sections should have been left as macros. It is difficult to assuring system wide inlining across modules.
102 remaining items
If I understand correctly, x86 official binaries are non-PGO builds.
Yeah, this is correct. We're more likely to deprecate and drop the 32-bit binaries before we make any major effort to optimise them - they run under an emulation layer in the OS (practically all supported OS installs are 64-bit native), so aren't really going to be recommended for people who care about performance anyway.
I think this issue can be closed. (I can't after migration)
Most of my experiences are invalid after Guido's #91718 corrected the quirks of MSVC.
Another reasonable fix would be a good test which makes specialized sections hotter.Thanks.
Closing as requested by OP. Thanks for your investigations @neonene ! Thanks to Guido too for the fix.
Reacted by neoneneThank you @neonene for your gentle pushes and encouragement and help to get this fixed!
Reacted by neoneneNo they are complementary.
Do you mean that this merged change 2f233fc is now useless?
No. What I said is about the optimization, not the (force) inlining. And what I suggested before have been already fixed by f8dc618 (and 2f233fc):
-
tp_*orcfuncpointer in the eval-loop can inline multiple callees without conflict. -
Moving
LOAD_FASTout of switch according to the scores below has no advantage now.
TOP3 entries with current 44 tests case 124 132522464 // LOAD_FAST case 100 48956231 // LOAD_CONST case 45 48318813 // LOAD_FAST__LOAD_FAST-
What I understand is that PGO build of Python 3.11 on Windows will be faster thanks to these changes, and the Windows python.org binaries only use PGO for 64-bit, not for 32-bit.
You can read a bit more posts and links because you have changed this thread's title several times.
Can someone please try to write a summary of this long and complex issue? It seems like different but related topics have been discussed and it's hard to get an overview. I'm confused between sometimes someone said that a change fixed the fix and then wrote that no, it's not really fixing the issue.
Reacted by Erlend E. AaslandLet me give it a quick try.
-
Originally, @neonene observed a Windows-specific performance regression in 3.10 between the a7 and b1 release. This was eventually shown to be caused by the function
_PyEval_EvalFrameDefaultgetting so long that the MSVC LTO gave up on inlining many things there. IIUC in 3.10 this was eventually fixed by making the function a bit smaller ([3.10] bpo-45116: Shrink interpreter by targetted revert of #25244 #28475). -
Of course, the same issue was then observed in the main branch (3.11). We then went back and forth trying various approaches to fix it. This wasn't easy because (a) the code kept changing (because the "Faster CPython" team was very active -- mostly growing the function), and (b) we had no good hardware or strategy to run reliable benchmarks. The latter problem spawned Measure Windows performance (and improve if lacking) faster-cpython/ideas#321.
-
Eventually I settled on a fix which consisted of turning a few inline functions back into macros, but only in ceval.c, and not in debug mode (and one only for MSVC). This was PR Performance regression 3.10b1: inlining issue in the big _PyEval_EvalFrameDefault() function with Visual Studio (MSC) #89279, commit 2f233fc.
-
Somewhat relatedly, I also figured out how to get MSVC to generate slightly faster switch code: if you switch on a one-byte value and all 256 cases exist, it skips a memory load. This was PR Improve performance for switch in ceval.c when using MSVC #91719, commit f8dc618.
-
Finally, I figured out how to get stable benchmark numbers (see Measure Windows performance (and improve if lacking) faster-cpython/ideas#321 (comment) and following comments) and showed that the macrofied inline functions gave us 10% performance back and the improved switch code gave 3%.
That's it.
-
Thanks for the summary. I would add that marking performance critical function with
__forceinline(Py_ALWAYS_INLINE) was tested, but it didn't work.
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: