Repository navigation
[C API] PEP 756: Add PyUnicode_Export() and PyUnicode_Import() functions #119609
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on May 27, 2024 - added a commit that references this issue
on May 27, 2024 cc @davidism
- added a commit that references this issue
on May 27, 2024 I believe we can do better. I'll write a longer reply this week.
Reacted by Steve DowerI believe we can do better.
My gut feeling as well. I don't think we want to encourage authors to specialise their own code when using the limited API. An efficient, abstraction-agnostic, multiple find-replace should be our job (if it's an important scenario, which I'm inclined to think it probably is).
My gut feeling as well. I don't think we want to encourage authors to specialise their own code when using the limited API.
C extensions already specialize their code for the Unicode native format (ex: MarkupSafe). The problem is that they cannot update their code to the limited API because of missing features to export to/import from native formats.
An efficient, abstraction-agnostic, multiple find-replace should be our job (if it's an important scenario, which I'm inclined to think it probably is).
That sounds like a very specific API solving only one use case.
C extensions already specialize their code for the Unicode native format
Which means we forced them into it, not that they wanted to. Just because we did something not great for them in the past means we have to double-down on it into the future.
That sounds like a very specific API solving only one use case.
The same could be said about your proposed API. Do you have more use cases?
I'm not going to argue heavily about this case until Petr provides his writeup. I haven't given it a huge amount of thought, just felt it was worth adding a +1 to "this doesn't feel great".
The same could be said about your proposed API.
What do you mean? I don't force users to use this API. It's an API which can be used to specialize code for UCS1, UCS2 and UCS4: it's the only way to write the most efficient code for the current implementation of Python.
Do you have more use cases?
Python itself has a wide library to specialize code for UCS1 / UCS2 / UCS4 in
Objects/stringlib/:Objects/stringlib/asciilib.h Objects/stringlib/codecs.h Objects/stringlib/count.h Objects/stringlib/ctype.h Objects/stringlib/eq.h Objects/stringlib/fastsearch.h Objects/stringlib/find.h Objects/stringlib/find_max_char.h Objects/stringlib/join.h Objects/stringlib/localeutil.h Objects/stringlib/partition.h Objects/stringlib/replace.h Objects/stringlib/split.h Objects/stringlib/stringdefs.h Objects/stringlib/transmogrify.h Objects/stringlib/ucs1lib.h Objects/stringlib/ucs2lib.h Objects/stringlib/ucs4lib.h Objects/stringlib/undef.h Objects/stringlib/unicode_format.hThere are different operations which are already optimized. The problem is that the macro PyUnicode_WRITE_CHAR() is implemented with 2 tests, it's inefficient. Well, in most cases, PyUnicode_WRITE_CHAR() is good enough. But not when you want the optimal case, such as MarkupSafe which has a single function optimized in C.
There are 27 projects using
PyUnicode_FromKindAndData()according to a code search in PyPI top 7,500 projects (2024-03-16):- Cython (3.0.9)
- Levenshtein (0.25.0)
- PyICU (2.12)
- PyICU-binary (2.7.4)
- PyQt5 (5.15.10)
- PyQt6 (6.6.1)
- aiocsv (1.3.1)
- asyncpg (0.29.0)
- biopython (1.83)
- catboost (1.2.3)
- cffi (1.16.0)
- mojimoji (0.0.13)
- mwparserfromhell (0.6.6)
- numba (0.59.0)
- numpy (1.26.4)
- orjson (3.9.15)
- pemja (0.4.1)
- pyahocorasick (2.0.0)
- pyjson5 (1.6.6)
- rapidfuzz (3.6.2)
- rcssmin (1.1.2)
- regex (2023.12.25)
- rjsmin (1.2.2)
- srsly (2.4.8)
- tokenizers (0.15.2)
- ujson (5.9.0)
- unicodedata2 (15.1.0)
There are already other, less efficient, ways to export a string, depending on its maximum character. But some of these functions are excluded from the limited C API.
- PyUnicode_AsUCS4() and PyUnicode_AsUCS4Copy() -- 4 bytes per character, it can waste memory
- PyUnicode_AsWideChar() and PyUnicode_AsWideCharString() -- need to handle surrogate pairs on Windows
- PyUnicode_AsUTF8String()
- PyUnicode_AsUTF16String()
- and all other encoding functions
Proposed
PyUnicode_AsNativeFormat()has a complexity of O(1) which is an important property, whereas these functions has at least a complexity of O(n) if not worse.Note: I don't understand why there is no PyUnicode_FromUCS4() function. So it's not possible to re-import a UCS4 string.Importing UCS4 can be done usingPyUnicode_FromKindAndData(PyUnicode_FromKindAndData).it's the only way to write the most efficient code for the current implementation of Python
"Most efficient code" is not the job of the limited API. At a certain level of exposure to internals, you need to stop using the limited API, or else accept that your performance will be reduced.
If the
AsUTF8/FromUTF8functions are not fast enough for markupsafe's multi-replace function, then as I said, I'd entertain a multi-replace function that doesn't leak implementation details, because that suits the limited API. I'd also consider some kind of builder API, which I believe you also have a proposal for.But this proposal is literally about leaking implementation details. That is against the intent of the limited API, and so I am against the proposal.
Thanks for making me think about it, you've upgraded me from "unsure" to "sure" ;) Still looking forward to Petr's thoughts.
@davidhewitt: Would Rust benefit from such "native format" API? Or does Rust just prefer UTF-8 and then decode UTF-8 in Python?
@da-woods: Would Cython benefit from such "native format" API? I see that Cython uses the following code in
__Pyx_PyUnicode_Substring():#if CYTHON_COMPILING_IN_LIMITED_API // PyUnicode_Substring() does not support negative indexing but is otherwise fine to use. return PyUnicode_Substring(text, start, stop); #else return PyUnicode_FromKindAndData(PyUnicode_KIND(text), PyUnicode_1BYTE_DATA(text) + start*PyUnicode_KIND(text), stop-start); #endif
At the very least, please rename it to "InternalFormat" instead of "NativeFormat".
At first I thought this was going to be a great API for following the conventions of the native machine, since the name led me incorrectly. It took me a few reads to figure out that it was telling me the format, not that I was getting to choose it.
I looked how some of these projects use
PyUnicode_FromKindAndData.- Cython uses
PyUnicode_Substringinstead ofPyUnicode_FromKindAndDatain the limited API.
https://github.com/cython/cython/blob/cf10ea12e3fc3637cf14e1bf962b295f53cc1a50/Cython/Utility/StringTools.c#L562-L568 - Levenshtein only uses
PyUnicode_FromKindAndDatawithPyUnicode_4BYTE_KIND. It could usePyUnicode_DecodeUTF32.
https://github.com/rapidfuzz/Levenshtein/blob/0c6d3dc4f3cdf80ce19392fd502733d80e6a5a5d/src/Levenshtein/levenshtein_cpp.pyx#L101 - PyQt5 only uses
PyUnicode_FromKindAndDatawithPyUnicode_2BYTE_KIND. It could usePyUnicode_DecodeUTF16. How does Qt5 represents non-BMP characters? If it uses UTF-16, thenPyUnicode_DecodeUTF16is the only correct solution.
https://github.com/baoboa/pyqt5/blob/11d5f43bc6f213d9d60272f3954a0048569cfc7c/designer/pluginloader.cpp#L174
https://github.com/baoboa/pyqt5/blob/11d5f43bc6f213d9d60272f3954a0048569cfc7c/qmlscene/pluginloader.cpp#L336 - aiocsv only uses
PyUnicode_FromKindAndDatawithPyUnicode_4BYTE_KIND.
https://github.com/MKuranowski/aiocsv/blob/0d8499a68cfe60b38d9f8fbab8c0b0c9f1b54c4a/aiocsv/_parser.c#L464 - asyncpg uses
PyUnicode_FromKindAndDatavia Cython either withPyUnicode_1BYTE_KINDfor short ASCII-only UUID strings (PyUnicode_FromString/PyUnicode_DecodeASCII/PyUnicode_DecodeLatin1/PyUnicode_DecodeUTF8can be used here) or withPyUnicode_4BYTE_KINDto create substrings from the buffer created byPyUnicode_AsUCS4Copy(PyUnicode_Substringcan be a little more efficient here).
https://github.com/lemoncode21/fastapi-reactjs-loginpage/blob/c156bf187d061030e221553b7bc94978839ec994/backend/venv/Lib/site-packages/asyncpg/pgproto/uuid.pyx#L168-L170
https://github.com/MIoCJluTeJllo/betwatchBot/blob/0c634eab1c47d6c421de2ab074d047658d2e362e/venv/Lib/site-packages/asyncpg/protocol/codecs/array.pyx#L644-L647
They usually use
PyUnicode_FromKindAndDatain only one or two places.PyUnicode_FromKindAndDataalways can be replaced with Latin1/UTF16/UTF32 decoder (with the "surrogatepass" error handler for UTF16/UTF32). I think this is why it was not included in the limited C API. The examined project do not need to usePyUnicode_FromKindAndData, there is always alternative in the limited C API.I do not see how
PyUnicode_AsNativeFormatandPyUnicode_FromNativeFormatcould help in the examined projects. They do not need conversion to/from internal format, they need conversion to/from a specific format. To utilize the benefit of working with the internal representation they need to reorganize their code in stringlib-like way, with a lot of macros and preprocessor directives (is it possible to do in Cython?) It has high implementation and maintenance cost. And it will likely will not work with PyUnicode_NATIVE_UTF8. And in some cases you need the UCS4 output buffer, because the stringlib-like way is bad for two-parameter specialization.Reacted by Steve Dower and Erlend E. Aasland- Cython uses
20 remaining items
I won't have time for a full review this week.
Ping. Do you have time for a review this week? :-)
Hi,
I sent my suggestions in a PR: vstinner#3
Let me know what you think :)Let me know what you think :)
There is an issue with your PR, it has something like 150 commits.
- added a commit that references this issue
on Jun 21, 2024 Yes, I merged in the main barnch to fix the conflicts. If you fix conflicts in your PR, I can rebase and just leave the last 6.
- Do we guarantee that the exported buffers zero-terminated? I assume we don't, to hide that implementation detail.
- IMO, 1-byte strings should be exportable to UCS2.
- I'd prefer renaming the size argument to nbytes to make the unit clear.
And of course I'd still prefer exporting
PyBuffer, rather than makingPyBuffer_Releasetake three arguments.@encukou: I integrated most of your suggestions in my PR. Please see the updated PR.
I created Add PyUnicode_Export() and PyUnicode_Import() to the limited C API issue in the C API WG Decisions project.
- changed the title
[-][C API] Add PyUnicode_Export() and PyUnicode_Import() functions[/-][+][C API] PEP 756: Add PyUnicode_Export() and PyUnicode_Import() functions[/+]on Sep 23, 2024 PEP 756 is withdrawn.
Feature or enhancement
PEP 393 – Flexible String Representation changed the Unicode implementation in Python 3.3 to use 3 string "kinds":
PyUnicode_KIND_1BYTE(UCS-1): ASCII and Latin1, [U+0000; U+00ff] range.PyUnicode_KIND_2BYTE(UCS-2): BMP, [U+0000; U+ffff] range.PyUnicode_KIND_4BYTE(UCZ-4): Full Unicode Character Set, [U+0000; U+10ffff] range.Strings must always use the optimal storage: ASCII string must be stored as PyUnicode_KIND_2BYTE.
Strings have a flag indicating if the string only contains ASCII characters: [U+0000; U+007f] range. It's used by multiple internal optimizations.
This implementation is not leaked in the limited C API. For example, the
PyUnicode_FromKindAndData()function is excluded from the stable ABI. Said differently, it's not possible to write efficient code for PEP 393 using the limited C API.I propose adding two functions:
PyUnicode_AsNativeFormat(): export to the native formatPyUnicode_FromNativeFormat(): import from the native formatThese functions are added to the limited C API version 3.14.
Native formats (new constants):
PyUnicode_NATIVE_ASCII: ASCII string.PyUnicode_NATIVE_UCS1: UCS-1 string.PyUnicode_NATIVE_UCS2: UCS-2 string.PyUnicode_NATIVE_UCS4: UCS-4 string.PyUnicode_NATIVE_UTF8: UTF-8 string (CPython implementation detail: only supported for import, not used by export).Differences with
PyUnicode_FromKindAndData():PyUnicode_NATIVE_ASCII format allows further optimizations.
PyUnicode_NATIVE_UTF8 can be used by PyPy and other Python implementation using UTF-8 as the internal storage.
API:
See the attached pull request for more details.
This feature was requested to me to port the MarkupSafe C extension to the limited C API. Currently, each release requires producing around 60 wheel files which takes 20 minutes to build: https://pypi.org/project/MarkupSafe/#files
Using the stable ABI would reduce the number of wheel packages and so ease their release process.
See src/markupsafe/_speedups.c: string functions specialized for the 3 string kinds (UCS-1, UCS-2, UCS-4).
Linked PRs