Repository navigation
BytesGenerator breaks UTF8 string #92081
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Apr 30, 2022 The problem has something to do with the maximum header length when the message is encoded to ASCII.
Increasing the maximum header length (the default is 78) makes the problem disappear for the sample tests:
g = email.generator.BytesGenerator(bytesmsg, maxheaderlen=160)Reacted by YuriThe problem has something to do with the maximum header length when the message is encoded to ASCII.
Increasing the maximum header length (the default is 78) makes the problem disappear for the sample tests:
g = email.generator.BytesGenerator(bytesmsg, maxheaderlen=160)Thanks for hint. As temporary workaround I set maxheaderlen to zero:
g = email.generator.BytesGenerator(bytesmsg, maxheaderlen=0)UPD.
For sending emails you should set max_line_length greater than 3. This is necessary in order to avoid ValueError at /Lib/email/quoprimime.py::body_encode() in some cases. I prefer to set max_line_length to maximum.msg = EmailMessage() msg.policy = msg.policy.clone(max_line_length=sys.maxsize)Note: This problem will occur with
Generatoras well asBytesGenerator.What I believe is happening in the problem cases is that the Generator is creating two or more encoded-words. A space character present in the input is placed between the encoded words rather than becoming a part of the encoded words. When the decoder is run, the decoder discards the linear whitespace that occurs between the encoded words. This leads to the omission of the space in the round tripped data.
Here's the output from the Generator for one of the problem cases:
Subject: =?utf-8?b?0YTRhNGE0YTRhNGE0YTRhNGE0YTRhNGE0YTRhNGE0YTRhNGE0YTRhNGE?= =?utf-8?b?0YQg0YQ=?=You can see that there are two spaces (in addition to the
\r\n) in between the two encoded blocks.After reading the examples in https://www.rfc-editor.org/rfc/rfc2047 I believe that the generator is at fault here. The space characters should be added to either the leading or trailing encoded-word rather than being emitted as a literal between the words.
Still exploring precisely what is going on but I have been able to modify the code to fix all of the cases reported by forcing whitespace to be encoded more frequently in
_header_value_parser.py::_refold_parse_tree()and preventing leading and trailing whitespace from being detected in_fold_as_ew(). I think my changes are too broad but I'll post them here in case something happens to me before I have a chance to finish diagnosing this:EDIT: Only leading whitespace needs to be undetected in
_fold_as_ew()diff --git a/Lib/email/_header_value_parser.py b/Lib/email/_header_value_parser.py index 8a8fb8bc42..498a3c1e01 100644 --- a/Lib/email/_header_value_parser.py +++ b/Lib/email/_header_value_parser.py @@ -2781,6 +2781,8 @@ def _refold_parse_tree(parse_tree, *, policy): if part.token_type == 'ptext' and set(tstr) & SPECIALS: # Encode if tstr contains special characters. want_encoding = True + elif part.token_type == 'fws' and last_ew: + want_encoding = True try: tstr.encode(encoding) charset = encoding @@ -2877,19 +2879,19 @@ def _fold_as_ew(to_encode, lines, maxlen, last_ew, ew_combine_allowed, charset): to_encode = str( get_unstructured(lines[-1][last_ew:] + to_encode)) lines[-1] = lines[-1][:last_ew] - if to_encode[0] in WSP: - # We're joining this to non-encoded text, so don't encode - # the leading blank. - leading_wsp = to_encode[0] - to_encode = to_encode[1:] - if (len(lines[-1]) == maxlen): - lines.append(_steal_trailing_WSP_if_exists(lines)) - lines[-1] += leading_wsp + #if to_encode[0] in WSP: + # # We're joining this to non-encoded text, so don't encode + # # the leading blank. + # leading_wsp = to_encode[0] + # to_encode = to_encode[1:] + # if (len(lines[-1]) == maxlen): + # lines.append(_steal_trailing_WSP_if_exists(lines)) + # lines[-1] += leading_wsp trailing_wsp = '' if to_encode[-1] in WSP: # Likewise for the trailing space.
Closer..... This change in
_fold_as_ewfixes all but two of the newly reported problems while all the unittests continue to pass:- if to_encode[0] in WSP: + if last_ew is None and to_encode[0] in WSP:
The two which are broken are:
ффффффффффффффффффффф ф фandф ффффффффффффффффффф ф фI believe something like the above should make it into the final fix as the comment for this block says that it should only be invoked if we're joining this encoded-word to a non-encoded word but there's nothing in this condition which checks that the previous word was non-encoded.
And
elif to_encode[0]:might also do the right thing here.- added a commit that references this issue
on May 4, 2022 I've opened a PR that fixes the cases where we need spaces between encoded words.
I've figured out that the remaining two problems start with a space. I'm searching the RFCs but so far haven't found anything that says whether the generator should handle this by including the initial space in the initial encoded word or if the decoder should handle it by displaying the space (or as a third option, that initial spaces in a Subject should be ignored).
I believe #92281 now fixes all the problems reported here. If anyone would like to test that the problems are resolved by it, that would be appreciated!
- added a commit that references this issue
on Jul 21, 2023 - added a commit that references this issue
on May 20, 2024 Thanks for the fix @abadger
- added a commit that references this issue
on Jul 17, 2024
Hi!
I found an issue when sending emails with Cyrillic letters in Subject header. Some spaces at Subject header are trimmed when sent.
Example:
When sending email with below subject:
Уведомление о принятии в работу обращенияat SMTP server logs I see subject that differs from original:
Уведомление о принятиив работу обращенияDuring research I've found that problem relates to small piece of code which encodes EmailMessage instance to byte string.
Python versions tested and problem confirmed: 3.8, 3.9, 3.10
Here is minimal reproducible example. Code can be used "as is", without any third party packages.
Minimal reproducible example
Above code demonstrates inequality of input and output strings after encoding message with BytesGenerator. Please note that not all strings with Cyrillic letters are broken. Only those strings that have word with single Cyrillic char only are affected under some conditions.
Small additional list of string with explanations you can find below:
Additional strings with explanations
Linked PRs