Repository navigation
test_unicode test_raiseMemError miscalculates struct size when string has UTF-8 representation. #93575
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errortestsTests in the Lib/test dirTests in the Lib/test dir3.11only security fixesonly security fixes3.12only security fixesonly security fixes
on Jun 7, 2022 - added a commit that references this issue
on Jun 7, 2022 No, it is not correct.
null_byteis the size of the null character in string "a". Does WASI use 2 or 4 bytes for character in the ASCII strings?It's a different issue.
sys.getsizeof()returns different values when one of the strings happened to have autf-8representation.Full test run:
0:15:27 Re-running test_unicode in verbose mode (matching: test_raiseMemError) test_raiseMemError (test.test_unicode.UnicodeTest.test_raiseMemError) ... test_raiseMemError (test.test_unicode.UnicodeTest.test_raiseMemError) (char='€', maxlen=1073741808, struct_size=31, char_size=2) ... FAIL ====================================================================== FAIL: test_raiseMemError (test.test_unicode.UnicodeTest.test_raiseMemError) (char='€', maxlen=1073741808, struct_size=31, char_size=2) ---------------------------------------------------------------------- Traceback (most recent call last): File "/Lib/test/test_unicode.py", line 2399, in test_raiseMemError self.assertRaises(MemoryError, alloc) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ AssertionError: MemoryError not raised by <lambda>Test run of just
test_unicodeERROR: test_raiseMemError (test.test_unicode.UnicodeTest.test_raiseMemError) (char='€', maxlen=1073741809, struct_size=28, char_size=2)struct_size=31!=struct_size=28. The extra 3 bytes arechar_size + null_byte. I'm going to change the test case.- changed the title
[-]test_unicode test_raiseMemError miscalculates length of NULL byte[/-][+]test_unicode test_raiseMemError miscalculates struct size when string has UTF-8 representation.[/+]on Jun 8, 2022 Ah, that's the thing! Nice catch.
I was going to use
struct_size = sys.getsizeof(char * n) - char_size * (n+1)
where
nis an arbitrary number.In case we add new fields the test will still pass successfully with your approach, but will no longer test the exact lower limit. Could you at least add a self-check for
sys.getsizeof(char * n) == struct_size + char_size * (n+1)?The assertion fails for
char='é':71 != 63(on 32 bit),99 != 83(on 64 bit)- added a commit that references this issue
on Jun 8, 2022 Because there is yet one bug here! The correct code for
struct_size:struct_size = ascii_struct_size if code < 0x80 else compact_struct_size
Yeah, the test was mishandling code between 128 and 256.
- added a commit that references this issue
on Jun 26, 2022
The test case
test_raiseMemErrorassumes that all structs have a NULL byte of length 1. Howevercompact_struct_sizeallocate 2 or 4 bytes space for NULL bytes:(PyUnicode_GET_LENGTH(self) + 1) * PyUnicode_KIND(self). Note that it's(len(s) + 1) * char_size, notlen(s) * char_size + 1. The bug introduces an off-by-one / off-by-three error that sometimes leads to failing test on WASI, because code does not raise a memory error.