Repository navigation
Add functions to get the width in columns of a character #56777
Description
Activity
Some characters take more than one column in a terminal, especially CJK (chinese, japanese, korean) characters. If you use such character in a terminal without taking care of the width in columns of each character, the text alignment can be broken. Issue bpo-2382 is an example of this problem.
bpo-2382 and bpo-6755 have patches implementing such function:
- unicode_width.patch of bpo-2382 adds unicode.width() method
- ucs2w.c of bpo-6755 creates a new ucs2w module with two functions: unichr2w() (width of a character) and ucs2w() (width of a string)
Use test_ucs2w.py of bpo-6755 to test these new functions/methods.
In the bpo-2382 code, how is the Windows case supposed to work? Also, what about systems that don't have wcswidth? IOW, the patch appears to be incorrect.
I like the bpo-6755 approach better, except that it shouldn't be using hard-coded tables, but instead integrate with Python's version of the UCD. In addition, it should use an accepted, published strategy for determining the width, preferably coming from the Unicode consortium.
I can attest that being able to get the columns of a grapheme cluster is very important for printing, because you need this to do correct linebreaking. There might be something you can steal from
http://search.cpan.org/perldoc?Unicode::GCString
http://search.cpan.org/perldoc?Unicode::LineBreakwhich implements UAX#14 on linebreaking and UAX#11 on East Asian widths.
I use this in my own code to help format Unicode strings my columns or lines. The right way would be to build this sort of knowledge into string.format(), but that is much harder, so an intermediary library module seems good enough for now.
There might be something you can steal from ...
I don't think that Python should reinvent the wheel. We should just reuse wcswidth().
Here is a simple patch exposing wcswidth() function as locale.width().
Example:
>>> import locale >>> text = '\u3042\u3044\u3046\u3048\u304a' >>> len(text) 5 >>> locale.width(text) 10 >>> locale.width(' ') 1 >>> locale.width('\U0010abcd') 1 >>> locale.width('\uDC80') Traceback (most recent call last): File "<stdin>", line 1, in <module> locale.Error: the string is not printable >>> locale.width('\U0010FFFF') Traceback (most recent call last): File "<stdin>", line 1, in <module> locale.Error: the string is not printable
I don't think that we need locale.width() on Windows because its console has already bigger issues with Unicode: see issue bpo-1602. If you want to display correctly non-ASCII characters on Windows, just avoid the Windows console and use a graphical widget.
Oh, unicode_width.patch of issue bpo-2382 implements the width on Windows using:
WideCharToMultiByte(CP_ACP, 0, buf, len, NULL, 0, NULL, NULL);
It computes the length of byte string encoded to the ANSI code page. I don't know if it can be seen as the "width" of a character string in the console...
I think the WideCharToMultibyte approach is just incorrect.
I'm -1 on using wcswidth, though. We already have unicodedata.east_asian_width, which implements http://unicode.org/reports/tr11/
The outcomes of this function are these:- F: full-width, width 2, compatibility character for a narrow char
- H: half-width, width 1, compatibility character for a narrow char
- W: wide, width 2
- Na: narrow, width 1
- A: ambiguous; width 2 in Asian context, width 1 in non-Asian context
- N: neutral; not used in Asian text, so has no width. Practically, width can be considered as 1
Martin v. Löwis <martin@v.loewis.de> added the comment:
I think the WideCharToMultibyte approach is just incorrect.
I'm -1 on using wcswidth, though.
Like you, I too seriously question using wcswidth() for this at all:
The wcswidth() function either shall return 0 (if pwcs points to a null wide-character code), or return the number of column positions to be occupied by the wide-character string pointed to by pwcs, or return -1 (if any of the first n wide-character codes in the wide- character string pointed to by pwcs is not a printable wide- character code).I would be willing to bet (a small amount of) money it does not correctly
inplmented Unicode print widths, even though one would certainly *think* it
does according to this:The wcswidth() function determines the number of column positions required for the first n characters of pwcs, or until a null wide character (L'\0') is encountered.There are a bunch of "interesting" cases I would want it tested against.
We already have unicodedata.east_asian_width, which implements http://unicode.org/reports/tr11/
The outcomes of this function are these:
- F: full-width, width 2, compatibility character for a narrow char
- H: half-width, width 1, compatibility character for a narrow char
- W: wide, width 2
- Na: narrow, width 1
- A: ambiguous; width 2 in Asian context, width 1 in non-Asian context
- N: neutral; not used in Asian text, so has no width. Practically, width can be considered as 1
Um, East_Asian_Width=Ambiguous (EA=A) isn't actually good enough for this.
And EA=N cannot be consider 1, either.For example, some of the Marks are EA=A and some are EA=N, yet how may
print columns they take varies. It is usually 0, but can be 1 at the start
of the file/string or immediately after a linebreak sequence. Then there
are things like the variation selectors which are never anything.Now consider the many \pC code points, like
U+0009 CHARACTER TABULATION U+00AD SOFT HYPHEN U+200C ZERO WIDTH NON-JOINER U+FEFF ZERO WIDTH NO-BREAK SPACE U+2062 INVISIBLE TIMESA TAB is its own problem but SHY we know is only width=1 immediately
before a linebreak or EOF, and ZWNJ and ZWNBSP are both certainly
width=0. So are the INVISIBLE * code points.Context:
Imagine you're trying to format a string so that it takes up exactly 20
columns: you need to know how many spaces to pad it with based on the
print width. That is what the bpo-12568 is needing
to do, and you have to do much more than East Asian Width properties.I really do think that what bpo-12568 is asking for is to have the equivalent
of the Perl Unicode::GCString's columns() method, and that you aren't going
to be able to handle text alignment of Unicode with anything that is much
less of that. After all, bpo-12568's title is "Add functions to get the width
in columns of a character". I would very much like to compare what
columns() thinks compared with what wcswidth() thinks. I bet wcswidth() is
very simple-minded at best.I may of course be wrong.
--tom
I'm -1 on using wcswidth, though.
When you write text into a console on Linux (e.g. displayed by gnome-terminal or konsole), I suppose that wcswidth() can be used to compute the width of a line. It would help to fix bpo-2382.
Or do you think that wcswidth() gives the wrong result for this use case?
> I'm -1 on using wcswidth, though.
When you write text into a console on Linux (e.g. displayed by
gnome-terminal or konsole), I suppose that wcswidth() can be used to
compute the width of a line. It would help to fix bpo-2382.Or do you think that wcswidth() gives the wrong result for this use
case?No, I think that using it is not necessary. If you want to compute the
width of a line, use unicodedata.east_asian_width. And yes, wcswidth
may sometimes produce "incorrect" results (although it's probably
correct most of the time).Could we have an update on the status of this? I ask because if 3.3 is going to (finally) fix unicode for curses, it would be really nice if it were possible to calculate the width of what's being displayed! It looks as if there was never quite agreement on the proper API....
Nicholas: I consider this issue fixed. There already *is* any API to compute the width of a character. Closing this as "works for me".
Martin: sorry to be completely dense, but I can't get this to work properly with the python3.3a1 build. Could you post some example code?
Please see the attached width.py for an example
30 remaining items
Hello,
I come from bpo-30717 . I have a pending PR that needs review ( #2673 ) adding a function that breaks unicode strings into grapheme clusters (aka what one would intuitively call "a character"). It's based on the grapheme cluster breaking algorithm from TR29.
Let me know if this is of any relevance.
Quick demo: >>> a=unicodedata.break_graphemes("lol") >>> list(a) ['l', 'o', 'l'] >>> list(unicodedata.break_graphemes("lo\u0309l")) ['l', 'ỏ', 'l'] >>> list(unicodedata.break_graphemes("lo\u0309\u0301l")) ['l', 'ỏ́', 'l'] >>> list(unicodedata.break_graphemes("lo\u0301l")) ['l', 'ó', 'l'] >>> list(unicodedata.break_graphemes("")) []
I suggest reclosing this issue, for the same reason I suggested closure of bpo-24665 in msg321291: abstract unicode 'characters' (graphemes) do not, in general, have fixed physical widths of 0, 1, or 2 n-pixel columns (or spaces). I based that fairly long message on IDLE's multiscript font sample as displayed on Windows 10. In that context, for instance, the width of (fixed-pitch) East Asian characters is about 1.6, not 2.0, times the width of fixed-pitch Ascii characters. Variable-width Tamil characters average about the same. The exact ratio depends on the Latin font used.
I did more experiments with Python started from Command Prompt with code page 437 or 65001 and characters 20 pixels high. The Windows console only allows 'fixed pitch' fonts. East Asian characters, if displayed, are expanded to double width.
However, European characters are not reliably displayed in one column. The width depends on the both the font selected when a character is entered and the current font. The 20 latin1 characters in '¢£¥§©«®¶½ĞÀÁÂÃÄÅÇÐØß' usually display in 20 columns. But if they are entered with the font set to MSGothic, the '§' and '¶' are each displayed in the middle of 2 columns, for 22 total. If the font is changed to MSGothic after entry, the '§' and '¶' are shifted 1/2 column right to overlap the following '©' or '½' without changing the total width. Greek and Cyrillic characters also sometimes take two columns.
I did not test whether the font size (pixel height) affects horizontal column spacing.
I close the issue as WONTFIX.
I think that even imperfect solution is better than no solution.
wcwidth()andwcswidth()are parts of Posix, they are used by modern terminals. There are several packages on PyPI:- https://pypi.org/project/wcwidth/
- https://pypi.org/project/cwcwidth/
- https://pypi.org/project/uwcwidth/
They are virtually mandatory for any Python program with rich CLI or TUI.
But now the stdlib itself needs such function for internal use -- in REPL,
calendar,argparse. We cannot depend on third-party packages.Reacted by Éric and Jeff Quast- added3.15bugs and security fixesbugs and security fixesand removed3.8 (EOL)end of lifeend of life
on Dec 8, 2025 But now the stdlib itself needs such function for internal use -- in REPL, calendar, argparse
On the "CJK support for textwrap issue", @methane wrote in 2018:
If someone really want this feature, please try it on PyPI.
And @merwok just closed the issue as "not planned": #68853
It seems like some core devs prefer to have a solution on PyPI rather than in the stdlib because of the complexity of the issue. Well, at least, the suggestion is to start the implementation on PyPI rather than starting in the stdlib.
I only changed the status from «closed as completed» to «closed as not planned» to better reflect the decision. But I support Serhiy’s position than something good enough is better than nothing, if the implementation is reasonable (i.e. using unicode data character classes if possible, not very long mappings like I saw in one of the third-party projects).
Interesting, that character with different from 1 is supported in formatting in C++23: https://timsong-cpp.github.io/cppwp/n4950/format.string.std
For example, while
format("{:*^6}", 'x')returns"**x***"(like"{:*^6}".format('x')in Python),format("{:*^6}", "🤡🤡🤡")returns"🤡🤡🤡", because the 🤡 character has width 2.So, we may end supporting this too, if even C++ supports this. Then
width()may be just a method ofstr, as Unicode aware 'isspace()',tolower()andsplitlines()which internally depend on the large Unicode database.Unfortunately, even if this is a part of the Posix and C++ standards, the exact behavior is not specified. I am looking at the code in glibc/g++, llvm and the Python's
wcwidth-- they are different, and even if they support the same rang (e.g. Hangul characters), they do it differently. We cannot have the implementation that matches the system behavior on Linux and macOS.Thus, it may be reason to provide also
locale.width(), like in the original patch. It should be used if we want absolute consistency with the system behavior.Reacted by Éric
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: