Skip to content

Add functions to get the width in columns of a character #56777

Description

@vstinner
BPO 12568
Nosy @malemburg, @loewis, @terryjreedy, @vstinner, @benjaminp, @ezio-melotti, @merwok, @bitdancer, @serhiy-storchaka, @Vermeille, @ishigoya, @bianjp
Files
  • locale_width.patch
  • width.py
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = <Date 2018-11-08.21:45:08.936>
    created_at = <Date 2011-07-14.22:43:56.356>
    labels = ['type-feature', '3.8', 'expert-unicode']
    title = 'Add functions to get the width in columns of a character'
    updated_at = <Date 2018-11-08.21:45:08.934>
    user = 'https://github.com/vstinner'

    bugs.python.org fields:

    activity = <Date 2018-11-08.21:45:08.934>
    actor = 'vstinner'
    assignee = 'none'
    closed = True
    closed_date = <Date 2018-11-08.21:45:08.936>
    closer = 'vstinner'
    components = ['Unicode']
    creation = <Date 2011-07-14.22:43:56.356>
    creator = 'vstinner'
    dependencies = []
    files = ['23401', '24773']
    hgrepos = []
    issue_num = 12568
    keywords = ['patch']
    message_count = 39.0
    messages = ['140376', '140488', '141936', '145497', '145498', '145523', '145535', '145748', '145778', '155223', '155236', '155307', '155313', '155323', '155324', '155337', '155342', '155343', '155344', '155345', '155346', '155361', '155370', '155373', '155379', '155382', '156337', '156348', '181149', '238425', '255421', '297129', '297489', '297492', '297564', '297569', '298322', '323731', '329488']
    nosy_count = 19.0
    nosy_names = ['lemburg', 'loewis', 'terry.reedy', 'vstinner', 'benjamin.peterson', 'ezio.melotti', 'eric.araujo', 'Arfrever', 'r.david.murray', 'inigoserna', 'zeha', 'poq', 'Nicholas.Cole', 'tchrist', 'serhiy.storchaka', 'Socob', 'Guillaume Sanchez', 'ishigoya', 'bianjp']
    pr_nums = []
    priority = 'normal'
    resolution = 'wont fix'
    stage = 'resolved'
    status = 'closed'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue12568'
    versions = ['Python 3.8']

    Activity

    1. vstinner commented on Jul 14, 2011

      @vstinner
      MemberAuthor

      Some characters take more than one column in a terminal, especially CJK (chinese, japanese, korean) characters. If you use such character in a terminal without taking care of the width in columns of each character, the text alignment can be broken. Issue bpo-2382 is an example of this problem.

      bpo-2382 and bpo-6755 have patches implementing such function:

      • unicode_width.patch of bpo-2382 adds unicode.width() method
      • ucs2w.c of bpo-6755 creates a new ucs2w module with two functions: unichr2w() (width of a character) and ucs2w() (width of a string)

      Use test_ucs2w.py of bpo-6755 to test these new functions/methods.

    2. loewis commented on Jul 16, 2011

      loewismannequin
      Mannequin

      In the bpo-2382 code, how is the Windows case supposed to work? Also, what about systems that don't have wcswidth? IOW, the patch appears to be incorrect.

      I like the bpo-6755 approach better, except that it shouldn't be using hard-coded tables, but instead integrate with Python's version of the UCD. In addition, it should use an accepted, published strategy for determining the width, preferably coming from the Unicode consortium.

    3. tchrist commented on Aug 12, 2011

      tchristmannequin
      Mannequin

      I can attest that being able to get the columns of a grapheme cluster is very important for printing, because you need this to do correct linebreaking. There might be something you can steal from

      http://search.cpan.org/perldoc?Unicode::GCString
      http://search.cpan.org/perldoc?Unicode::LineBreak

      which implements UAX#14 on linebreaking and UAX#11 on East Asian widths.

      I use this in my own code to help format Unicode strings my columns or lines. The right way would be to build this sort of knowledge into string.format(), but that is much harder, so an intermediary library module seems good enough for now.

    4. vstinner commented on Oct 14, 2011

      @vstinner
      MemberAuthor

      There might be something you can steal from ...

      I don't think that Python should reinvent the wheel. We should just reuse wcswidth().

      Here is a simple patch exposing wcswidth() function as locale.width().

      Example:

      >>> import locale
      >>> text = '\u3042\u3044\u3046\u3048\u304a'
      >>> len(text)
      5
      >>> locale.width(text)
      10
      >>> locale.width(' ')
      1
      >>> locale.width('\U0010abcd')
      1
      >>> locale.width('\uDC80')
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      locale.Error: the string is not printable
      >>> locale.width('\U0010FFFF')
      Traceback (most recent call last):
        File "<stdin>", line 1, in <module>
      locale.Error: the string is not printable

      I don't think that we need locale.width() on Windows because its console has already bigger issues with Unicode: see issue bpo-1602. If you want to display correctly non-ASCII characters on Windows, just avoid the Windows console and use a graphical widget.

    5. vstinner commented on Oct 14, 2011

      @vstinner
      MemberAuthor

      Oh, unicode_width.patch of issue bpo-2382 implements the width on Windows using:

      WideCharToMultiByte(CP_ACP, 0, buf, len, NULL, 0, NULL, NULL);

      It computes the length of byte string encoded to the ANSI code page. I don't know if it can be seen as the "width" of a character string in the console...

    6. loewis commented on Oct 14, 2011

      loewismannequin
      Mannequin

      I think the WideCharToMultibyte approach is just incorrect.

      I'm -1 on using wcswidth, though. We already have unicodedata.east_asian_width, which implements http://unicode.org/reports/tr11/
      The outcomes of this function are these:

      • F: full-width, width 2, compatibility character for a narrow char
      • H: half-width, width 1, compatibility character for a narrow char
      • W: wide, width 2
      • Na: narrow, width 1
      • A: ambiguous; width 2 in Asian context, width 1 in non-Asian context
      • N: neutral; not used in Asian text, so has no width. Practically, width can be considered as 1
    7. tchrist commented on Oct 14, 2011

      tchristmannequin
      Mannequin

      Martin v. Löwis <martin@v.loewis.de> added the comment:

      I think the WideCharToMultibyte approach is just incorrect.

      I'm -1 on using wcswidth, though.

      Like you, I too seriously question using wcswidth() for this at all:

      The wcswidth() function either shall return 0 (if pwcs points to a
      null wide-character code), or return the number of column positions
      to be occupied by the wide-character string pointed to by pwcs, or
      return -1 (if any of the first n wide-character codes in the wide-
      character string pointed to by pwcs is not a printable wide-
      character code).
      

      I would be willing to bet (a small amount of) money it does not correctly
      inplmented Unicode print widths, even though one would certainly *think* it
      does according to this:

       The wcswidth() function determines the number of column positions
       required for the first n characters of pwcs, or until a null wide
       character (L'\0') is encountered.
      

      There are a bunch of "interesting" cases I would want it tested against.

      We already have unicodedata.east_asian_width, which implements http://unicode.org/reports/tr11/

      The outcomes of this function are these:

      • F: full-width, width 2, compatibility character for a narrow char
      • H: half-width, width 1, compatibility character for a narrow char
      • W: wide, width 2
      • Na: narrow, width 1
      • A: ambiguous; width 2 in Asian context, width 1 in non-Asian context
      • N: neutral; not used in Asian text, so has no width. Practically, width can be considered as 1

      Um, East_Asian_Width=Ambiguous (EA=A) isn't actually good enough for this.
      And EA=N cannot be consider 1, either.

      For example, some of the Marks are EA=A and some are EA=N, yet how may
      print columns they take varies. It is usually 0, but can be 1 at the start
      of the file/string or immediately after a linebreak sequence. Then there
      are things like the variation selectors which are never anything.

      Now consider the many \pC code points, like

      U+0009  CHARACTER TABULATION
      U+00AD  SOFT HYPHEN 
      U+200C  ZERO WIDTH NON-JOINER
      U+FEFF  ZERO WIDTH NO-BREAK SPACE
      U+2062  INVISIBLE TIMES
      

      A TAB is its own problem but SHY we know is only width=1 immediately
      before a linebreak or EOF, and ZWNJ and ZWNBSP are both certainly
      width=0. So are the INVISIBLE * code points.

      Context:

      Imagine you're trying to format a string so that it takes up exactly 20
      columns: you need to know how many spaces to pad it with based on the
      print width. That is what the bpo-12568 is needing
      to do, and you have to do much more than East Asian Width properties.

      I really do think that what bpo-12568 is asking for is to have the equivalent
      of the Perl Unicode::GCString's columns() method, and that you aren't going
      to be able to handle text alignment of Unicode with anything that is much
      less of that. After all, bpo-12568's title is "Add functions to get the width
      in columns of a character". I would very much like to compare what
      columns() thinks compared with what wcswidth() thinks. I bet wcswidth() is
      very simple-minded at best.

      I may of course be wrong.

      --tom

    8. vstinner commented on Oct 17, 2011

      @vstinner
      MemberAuthor

      I'm -1 on using wcswidth, though.

      When you write text into a console on Linux (e.g. displayed by gnome-terminal or konsole), I suppose that wcswidth() can be used to compute the width of a line. It would help to fix bpo-2382.

      Or do you think that wcswidth() gives the wrong result for this use case?

    9. loewis commented on Oct 18, 2011

      loewismannequin
      Mannequin

      > I'm -1 on using wcswidth, though.

      When you write text into a console on Linux (e.g. displayed by
      gnome-terminal or konsole), I suppose that wcswidth() can be used to
      compute the width of a line. It would help to fix bpo-2382.

      Or do you think that wcswidth() gives the wrong result for this use
      case?

      No, I think that using it is not necessary. If you want to compute the
      width of a line, use unicodedata.east_asian_width. And yes, wcswidth
      may sometimes produce "incorrect" results (although it's probably
      correct most of the time).

    10. NicholasCole commented on Mar 9, 2012

      NicholasColemannequin
      Mannequin

      Could we have an update on the status of this? I ask because if 3.3 is going to (finally) fix unicode for curses, it would be really nice if it were possible to calculate the width of what's being displayed! It looks as if there was never quite agreement on the proper API....

    11. loewis commented on Mar 9, 2012

      loewismannequin
      Mannequin

      Nicholas: I consider this issue fixed. There already *is* any API to compute the width of a character. Closing this as "works for me".

    12. NicholasCole commented on Mar 10, 2012

      NicholasColemannequin
      Mannequin

      Martin: sorry to be completely dense, but I can't get this to work properly with the python3.3a1 build. Could you post some example code?

    13. loewis commented on Mar 10, 2012

      loewismannequin
      Mannequin

      Please see the attached width.py for an example

    14. 30 remaining items

    15. Vermeille commented on Jul 13, 2017

      Vermeillemannequin
      Mannequin

      Hello,

      I come from bpo-30717 . I have a pending PR that needs review ( #2673 ) adding a function that breaks unicode strings into grapheme clusters (aka what one would intuitively call "a character"). It's based on the grapheme cluster breaking algorithm from TR29.

      Let me know if this is of any relevance.

      Quick demo:
      >>> a=unicodedata.break_graphemes("lol")
      >>> list(a)
      ['l', 'o', 'l']
      >>> list(unicodedata.break_graphemes("lo\u0309l"))
      ['l', 'ỏ', 'l']
      >>> list(unicodedata.break_graphemes("lo\u0309\u0301l"))
      ['l', 'ỏ́', 'l']
      >>> list(unicodedata.break_graphemes("lo\u0301l"))
      ['l', 'ó', 'l']
      >>> list(unicodedata.break_graphemes(""))
      []
    16. terryjreedy commented on Aug 18, 2018

      @terryjreedy
      Member

      I suggest reclosing this issue, for the same reason I suggested closure of bpo-24665 in msg321291: abstract unicode 'characters' (graphemes) do not, in general, have fixed physical widths of 0, 1, or 2 n-pixel columns (or spaces). I based that fairly long message on IDLE's multiscript font sample as displayed on Windows 10. In that context, for instance, the width of (fixed-pitch) East Asian characters is about 1.6, not 2.0, times the width of fixed-pitch Ascii characters. Variable-width Tamil characters average about the same. The exact ratio depends on the Latin font used.

      I did more experiments with Python started from Command Prompt with code page 437 or 65001 and characters 20 pixels high. The Windows console only allows 'fixed pitch' fonts. East Asian characters, if displayed, are expanded to double width.

      However, European characters are not reliably displayed in one column. The width depends on the both the font selected when a character is entered and the current font. The 20 latin1 characters in '¢£¥§©«®¶½ĞÀÁÂÃÄÅÇÐØß' usually display in 20 columns. But if they are entered with the font set to MSGothic, the '§' and '¶' are each displayed in the middle of 2 columns, for 22 total. If the font is changed to MSGothic after entry, the '§' and '¶' are shifted 1/2 column right to overlap the following '©' or '½' without changing the total width. Greek and Cyrillic characters also sometimes take two columns.

      I did not test whether the font size (pixel height) affects horizontal column spacing.

    17. vstinner commented on Nov 8, 2018

      @vstinner
      MemberAuthor

      I close the issue as WONTFIX.

    18. transferred this issue fromon Apr 10, 2022
    19. serhiy-storchaka commented on Dec 8, 2025

      @serhiy-storchaka
      Member

      I think that even imperfect solution is better than no solution. wcwidth() and wcswidth() are parts of Posix, they are used by modern terminals. There are several packages on PyPI:

      They are virtually mandatory for any Python program with rich CLI or TUI.

      But now the stdlib itself needs such function for internal use -- in REPL, calendar, argparse. We cannot depend on third-party packages.

    20. added
      3.15bugs and security fixes
      and removed on Dec 8, 2025
    21. vstinner commented on Dec 9, 2025

      @vstinner
      MemberAuthor

      But now the stdlib itself needs such function for internal use -- in REPL, calendar, argparse

      On the "CJK support for textwrap issue", @methane wrote in 2018:

      If someone really want this feature, please try it on PyPI.

      And @merwok just closed the issue as "not planned": #68853

      It seems like some core devs prefer to have a solution on PyPI rather than in the stdlib because of the complexity of the issue. Well, at least, the suggestion is to start the implementation on PyPI rather than starting in the stdlib.

    22. merwok commented on Dec 9, 2025

      @merwok
      Member

      I only changed the status from «closed as completed» to «closed as not planned» to better reflect the decision. But I support Serhiy’s position than something good enough is better than nothing, if the implementation is reasonable (i.e. using unicode data character classes if possible, not very long mappings like I saw in one of the third-party projects).

    23. serhiy-storchaka commented on Dec 10, 2025

      @serhiy-storchaka
      Member

      Interesting, that character with different from 1 is supported in formatting in C++23: https://timsong-cpp.github.io/cppwp/n4950/format.string.std

      For example, while format("{:*^6}", 'x') returns "**x***" (like "{:*^6}".format('x') in Python), format("{:*^6}", "🤡🤡🤡") returns "🤡🤡🤡", because the 🤡 character has width 2.

      So, we may end supporting this too, if even C++ supports this. Then width() may be just a method of str, as Unicode aware 'isspace()', tolower() and splitlines() which internally depend on the large Unicode database.

      Unfortunately, even if this is a part of the Posix and C++ standards, the exact behavior is not specified. I am looking at the code in glibc/g++, llvm and the Python's wcwidth -- they are different, and even if they support the same rang (e.g. Hangul characters), they do it differently. We cannot have the implementation that matches the system behavior on Linux and macOS.

      Thus, it may be reason to provide also locale.width(), like in the original patch. It should be used if we want absolute consistency with the system behavior.

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions