Skip to content

[C API] PEP 756: Add PyUnicode_Export() and PyUnicode_Import() functions #119609

Description

@vstinner

Feature or enhancement

PEP 393 – Flexible String Representation changed the Unicode implementation in Python 3.3 to use 3 string "kinds":

  • PyUnicode_KIND_1BYTE (UCS-1): ASCII and Latin1, [U+0000; U+00ff] range.
  • PyUnicode_KIND_2BYTE (UCS-2): BMP, [U+0000; U+ffff] range.
  • PyUnicode_KIND_4BYTE (UCZ-4): Full Unicode Character Set, [U+0000; U+10ffff] range.

Strings must always use the optimal storage: ASCII string must be stored as PyUnicode_KIND_2BYTE.

Strings have a flag indicating if the string only contains ASCII characters: [U+0000; U+007f] range. It's used by multiple internal optimizations.

This implementation is not leaked in the limited C API. For example, the PyUnicode_FromKindAndData() function is excluded from the stable ABI. Said differently, it's not possible to write efficient code for PEP 393 using the limited C API.


I propose adding two functions:

  • PyUnicode_AsNativeFormat(): export to the native format
  • PyUnicode_FromNativeFormat(): import from the native format

These functions are added to the limited C API version 3.14.

Native formats (new constants):

  • PyUnicode_NATIVE_ASCII: ASCII string.
  • PyUnicode_NATIVE_UCS1: UCS-1 string.
  • PyUnicode_NATIVE_UCS2: UCS-2 string.
  • PyUnicode_NATIVE_UCS4: UCS-4 string.
  • PyUnicode_NATIVE_UTF8: UTF-8 string (CPython implementation detail: only supported for import, not used by export).

Differences with PyUnicode_FromKindAndData():

  • Size is a number of bytes. For example, a single UCS-2 character is counted as 2 bytes.
  • Add PyUnicode_NATIVE_ASCII and PyUnicode_NATIVE_UTF8 formats.

PyUnicode_NATIVE_ASCII format allows further optimizations.

PyUnicode_NATIVE_UTF8 can be used by PyPy and other Python implementation using UTF-8 as the internal storage.


API:

#define PyUnicode_NATIVE_ASCII 1
#define PyUnicode_NATIVE_UCS1 2
#define PyUnicode_NATIVE_UCS2 3
#define PyUnicode_NATIVE_UCS4 4
#define PyUnicode_NATIVE_UTF8 5

// Get the content of a string in its native format.
// - Return the content, set '*size' and '*native_format' on success.
// - Set an exception and return NULL on error.
PyAPI_FUNC(const void*) PyUnicode_AsNativeFormat(
    PyObject *unicode,
    Py_ssize_t *size,
    int *native_format);

// Create a string object from a native format string.
// - Return a reference to a new string object on success.
// - Set an exception and return NULL on error.
PyAPI_FUNC(PyObject*) PyUnicode_FromNativeFormat(
    const void *data,
    Py_ssize_t size,
    int native_format);

See the attached pull request for more details.


This feature was requested to me to port the MarkupSafe C extension to the limited C API. Currently, each release requires producing around 60 wheel files which takes 20 minutes to build: https://pypi.org/project/MarkupSafe/#files

Using the stable ABI would reduce the number of wheel packages and so ease their release process.

See src/markupsafe/_speedups.c: string functions specialized for the 3 string kinds (UCS-1, UCS-2, UCS-4).

Linked PRs

Activity

  1. added a commit that references this issue on May 27, 2024
  2. vstinner commented on May 27, 2024

    @vstinner
    MemberAuthor
  3. vstinner commented on May 27, 2024

    @vstinner
    MemberAuthor
  4. added a commit that references this issue on May 27, 2024
  5. encukou commented on May 28, 2024

    @encukou
    Member

    I believe we can do better. I'll write a longer reply this week.

  6. zooba commented on May 28, 2024

    @zooba
    Member

    I believe we can do better.

    My gut feeling as well. I don't think we want to encourage authors to specialise their own code when using the limited API. An efficient, abstraction-agnostic, multiple find-replace should be our job (if it's an important scenario, which I'm inclined to think it probably is).

  7. vstinner commented on May 28, 2024

    @vstinner
    MemberAuthor

    My gut feeling as well. I don't think we want to encourage authors to specialise their own code when using the limited API.

    C extensions already specialize their code for the Unicode native format (ex: MarkupSafe). The problem is that they cannot update their code to the limited API because of missing features to export to/import from native formats.

    An efficient, abstraction-agnostic, multiple find-replace should be our job (if it's an important scenario, which I'm inclined to think it probably is).

    That sounds like a very specific API solving only one use case.

  8. zooba commented on May 28, 2024

    @zooba
    Member

    C extensions already specialize their code for the Unicode native format

    Which means we forced them into it, not that they wanted to. Just because we did something not great for them in the past means we have to double-down on it into the future.

    That sounds like a very specific API solving only one use case.

    The same could be said about your proposed API. Do you have more use cases?

    I'm not going to argue heavily about this case until Petr provides his writeup. I haven't given it a huge amount of thought, just felt it was worth adding a +1 to "this doesn't feel great".

  9. vstinner commented on May 28, 2024

    @vstinner
    MemberAuthor

    The same could be said about your proposed API.

    What do you mean? I don't force users to use this API. It's an API which can be used to specialize code for UCS1, UCS2 and UCS4: it's the only way to write the most efficient code for the current implementation of Python.

    Do you have more use cases?

    Python itself has a wide library to specialize code for UCS1 / UCS2 / UCS4 in Objects/stringlib/:

    Objects/stringlib/asciilib.h
    Objects/stringlib/codecs.h
    Objects/stringlib/count.h
    Objects/stringlib/ctype.h
    Objects/stringlib/eq.h
    Objects/stringlib/fastsearch.h
    Objects/stringlib/find.h
    Objects/stringlib/find_max_char.h
    Objects/stringlib/join.h
    Objects/stringlib/localeutil.h
    Objects/stringlib/partition.h
    Objects/stringlib/replace.h
    Objects/stringlib/split.h
    Objects/stringlib/stringdefs.h
    Objects/stringlib/transmogrify.h
    Objects/stringlib/ucs1lib.h
    Objects/stringlib/ucs2lib.h
    Objects/stringlib/ucs4lib.h
    Objects/stringlib/undef.h
    Objects/stringlib/unicode_format.h
    

    There are different operations which are already optimized. The problem is that the macro PyUnicode_WRITE_CHAR() is implemented with 2 tests, it's inefficient. Well, in most cases, PyUnicode_WRITE_CHAR() is good enough. But not when you want the optimal case, such as MarkupSafe which has a single function optimized in C.

    There are 27 projects using PyUnicode_FromKindAndData() according to a code search in PyPI top 7,500 projects (2024-03-16):

    • Cython (3.0.9)
    • Levenshtein (0.25.0)
    • PyICU (2.12)
    • PyICU-binary (2.7.4)
    • PyQt5 (5.15.10)
    • PyQt6 (6.6.1)
    • aiocsv (1.3.1)
    • asyncpg (0.29.0)
    • biopython (1.83)
    • catboost (1.2.3)
    • cffi (1.16.0)
    • mojimoji (0.0.13)
    • mwparserfromhell (0.6.6)
    • numba (0.59.0)
    • numpy (1.26.4)
    • orjson (3.9.15)
    • pemja (0.4.1)
    • pyahocorasick (2.0.0)
    • pyjson5 (1.6.6)
    • rapidfuzz (3.6.2)
    • rcssmin (1.1.2)
    • regex (2023.12.25)
    • rjsmin (1.2.2)
    • srsly (2.4.8)
    • tokenizers (0.15.2)
    • ujson (5.9.0)
    • unicodedata2 (15.1.0)

    There are already other, less efficient, ways to export a string, depending on its maximum character. But some of these functions are excluded from the limited C API.

    • PyUnicode_AsUCS4() and PyUnicode_AsUCS4Copy() -- 4 bytes per character, it can waste memory
    • PyUnicode_AsWideChar() and PyUnicode_AsWideCharString() -- need to handle surrogate pairs on Windows
    • PyUnicode_AsUTF8String()
    • PyUnicode_AsUTF16String()
    • and all other encoding functions

    Proposed PyUnicode_AsNativeFormat() has a complexity of O(1) which is an important property, whereas these functions has at least a complexity of O(n) if not worse.

    Note: I don't understand why there is no PyUnicode_FromUCS4() function. So it's not possible to re-import a UCS4 string. Importing UCS4 can be done using PyUnicode_FromKindAndData(PyUnicode_FromKindAndData).

  10. zooba commented on May 28, 2024

    @zooba
    Member

    it's the only way to write the most efficient code for the current implementation of Python

    "Most efficient code" is not the job of the limited API. At a certain level of exposure to internals, you need to stop using the limited API, or else accept that your performance will be reduced.

    If the AsUTF8/FromUTF8 functions are not fast enough for markupsafe's multi-replace function, then as I said, I'd entertain a multi-replace function that doesn't leak implementation details, because that suits the limited API. I'd also consider some kind of builder API, which I believe you also have a proposal for.

    But this proposal is literally about leaking implementation details. That is against the intent of the limited API, and so I am against the proposal.

    Thanks for making me think about it, you've upgraded me from "unsure" to "sure" ;) Still looking forward to Petr's thoughts.

  11. vstinner commented on May 28, 2024

    @vstinner
    MemberAuthor

    @davidhewitt: Would Rust benefit from such "native format" API? Or does Rust just prefer UTF-8 and then decode UTF-8 in Python?

    @da-woods: Would Cython benefit from such "native format" API? I see that Cython uses the following code in __Pyx_PyUnicode_Substring():

    #if CYTHON_COMPILING_IN_LIMITED_API
        // PyUnicode_Substring() does not support negative indexing but is otherwise fine to use.
        return PyUnicode_Substring(text, start, stop);
    #else
        return PyUnicode_FromKindAndData(PyUnicode_KIND(text),
            PyUnicode_1BYTE_DATA(text) + start*PyUnicode_KIND(text), stop-start);
    #endif
  12. zooba commented on May 28, 2024

    @zooba
    Member

    At the very least, please rename it to "InternalFormat" instead of "NativeFormat".

    At first I thought this was going to be a great API for following the conventions of the native machine, since the name led me incorrectly. It took me a few reads to figure out that it was telling me the format, not that I was getting to choose it.

  13. serhiy-storchaka commented on May 28, 2024

    @serhiy-storchaka
    Member

    I looked how some of these projects use PyUnicode_FromKindAndData.

    They usually use PyUnicode_FromKindAndData in only one or two places. PyUnicode_FromKindAndData always can be replaced with Latin1/UTF16/UTF32 decoder (with the "surrogatepass" error handler for UTF16/UTF32). I think this is why it was not included in the limited C API. The examined project do not need to use PyUnicode_FromKindAndData, there is always alternative in the limited C API.

    I do not see how PyUnicode_AsNativeFormat and PyUnicode_FromNativeFormat could help in the examined projects. They do not need conversion to/from internal format, they need conversion to/from a specific format. To utilize the benefit of working with the internal representation they need to reorganize their code in stringlib-like way, with a lot of macros and preprocessor directives (is it possible to do in Cython?) It has high implementation and maintenance cost. And it will likely will not work with PyUnicode_NATIVE_UTF8. And in some cases you need the UCS4 output buffer, because the stringlib-like way is bad for two-parameter specialization.

  14. 20 remaining items

  15. vstinner commented on Jun 19, 2024

    @vstinner
    MemberAuthor

    @encukou:

    I won't have time for a full review this week.

    Ping. Do you have time for a review this week? :-)

  16. encukou commented on Jun 20, 2024

    @encukou
    Member

    Hi,
    I sent my suggestions in a PR: vstinner#3
    Let me know what you think :)

  17. vstinner commented on Jun 20, 2024

    @vstinner
    MemberAuthor

    Let me know what you think :)

    There is an issue with your PR, it has something like 150 commits.

  18. added a commit that references this issue on Jun 21, 2024
  19. encukou commented on Jun 21, 2024

    @encukou
    Member

    Yes, I merged in the main barnch to fix the conflicts. If you fix conflicts in your PR, I can rebase and just leave the last 6.

    • Do we guarantee that the exported buffers zero-terminated? I assume we don't, to hide that implementation detail.
    • IMO, 1-byte strings should be exportable to UCS2.
    • I'd prefer renaming the size argument to nbytes to make the unit clear.

    And of course I'd still prefer exporting PyBuffer, rather than making PyBuffer_Release take three arguments.

  20. vstinner commented on Jun 21, 2024

    @vstinner
    MemberAuthor

    @encukou: I integrated most of your suggestions in my PR. Please see the updated PR.

  21. vstinner commented on Jun 24, 2024

    @vstinner
    MemberAuthor

    I created Add PyUnicode_Export() and PyUnicode_Import() to the limited C API issue in the C API WG Decisions project.

  22. added 4 commits that reference this issue on Sep 5, 2024
  23. changed the title [-][C API] Add PyUnicode_Export() and PyUnicode_Import() functions[/-] [+][C API] PEP 756: Add PyUnicode_Export() and PyUnicode_Import() functions[/+] on Sep 23, 2024
  24. vstinner commented on Dec 6, 2024

    @vstinner
    MemberAuthor

    PEP 756 is withdrawn.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions