Repository navigation
"z" format specifier is treated differently in unicode and bytes #104018
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Apr 30, 2023 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Apr 30, 2023 Good catch. I think having that enabled was left over from an early incarnation of the PR, before it was decided that %-format would not be supported.
For now, I confirmed that the tests still pass after disable of
case 'z'in_PyBytes_FormatEx().Fix should include tests to confirm that "z" format is not accepted for %-formatting.
- added a commit that references this issue
on May 1, 2023 - added a commit that references this issue
on May 1, 2023 - added a commit that references this issue
on May 1, 2023 @belm0, @mdickinson, thanks for the prompt fix.
May I ask why
static formatfloat()in bytesobject.c remains withF_NO_NEG_0handling? Offhand it looks like that flag bit could never make it into that function, but I might be missing something. The same question applies to unicodeobject.c .Reacted by John BelmonteMay I ask why
static formatfloat()in bytesobject.c remains withF_NO_NEG_0handling? Offhand it looks like that flag bit could never make it into that function, but I might be missing something. The same question applies to unicodeobject.c .Thank you, please see #104107
Thanks
- added a commit that references this issue
on May 7, 2023 - added a commit that references this issue
on May 7, 2023
Hello up there. I've hit a discrepancy in how
zflags is handled by%in unicode and bytes:%rejects it as "unsupported format character" according to original discussion in string formatting: normalize negative zero #90153 (= BPO-45995),%fully handles "z":In other words there is inconsistency in how 'z' is handled by '%' for unicode and bytes, and there is also inconsistency in how 'z' was supposed to be handled by
.formatand not handled by '%' as originally discussed on BPO-45995.'z' handling was implemented in #30049 and indeed there I see b'%z' being fully handled:
b0b836b20cb5#diff-f6d440aad34e1c4535c0d898c0197a95490766c745991caace6f64b5dd1ece51
but u'%z' being only partly handled internally without corresponding frontend parsing that bytes has:
b0b836b20cb5#diff-34c966e7876d6f8bf801dd51896327e4f68bba02cddb95fbf3963f0b2e39c38a
In my view the fix should be either a) to add '%z' handling to unicode, or b) to remove '%z' handling from bytes.
Thanks beforehand,
Kirill
/cc @belm0, @mdickinson
Linked PRs