Repository navigation
Python 3.9.14: grammar/clarity improvements for str.encode, str.decode error-checking documentation #99991
Description
Activity
do you have recommendations on how to improve them
It's not easy, but I'll try - here's an attempt for
str.encode:str.encode(encoding="utf-8", errors="strict")Return a bytes object containing the string in the requested encoding.
The default
stricterror handling raises a UnicodeError exception when an encoding error occurs. Other possible values includeignore,replace,xmlcharrefreplace,backslashreplaceand any other name registered via codecs.register_error(), see section Error Handlers.For a list of possible encodings, see section Standard Encodings.
For performance reasons, the value of the errors argument is not checked for validity unless Python Development Mode is enabled.
Process used
Thinking like refactoring, the important items to maintain (like creating unit tests) seemed to be:- That the documentation describes what the method does (encodes a string, handles errors)
- That relevant entities are hyperlinked/emphasized
- That the documentation is relatively concise (the time required to read and understand the description is similar or reduced)
So looking at the text, there seemed to be a few steps to refactor it:
- Break the dense paragraph into individual statements (usually individual sentences)
- Apply consistent emphasis to entities (types, arguments, ...) that appear in the statements
- Remove redundant/self-explanatory statements
- It's a 2-argument method -- and each argument requires explanation -- so consolidate the statements into one paragraph per argument:
- Move sentences describing
encodinginto the first paragraph - Move sentences describing
errorsinto the second paragraph - Exception: some of the commentary appears to be about non-functional behaviour (performance) -- OK, that could move into a third paragraph
- Move sentences describing
- Re-read and refactor to remove unnecessary words and statements
- Nitpick: error handling is important to understand, so let's move the list of codecs -- useful though it is -- to after the errors argument, in the hope that people will read about error-handling first
Note:
- https://docs.python.org/3.9/library/codecs.html#codecs.EncodedFile also has documentation for a similarly-behaved
errorsargument
Thanks. I'll work on that.
Reacted by James AddisonFor performance reasons, the value of the errors argument is not checked for validity unless Python Development Mode is enabled.
Ah: I forgot to mention debug builds in this suggestion. Similar to development-mode, debug builds also validate the errors argument.
A simple check to test whether error-handler-name-validation is enabled is to attempt encoding using an invalid error-handler-name. For example,
"".encode('utf-8', errors='gh-99991').By the way, thanks for your comprehensive and insightful feedback, @jayaddison ! We'd love to have more of it elsewhere, and you're welcome to contribute the changes yourself if you like—you seem to have quite a knack for documentation writing, which makes a big positive impact on the community. Thanks again!
Reacted by James AddisonThanks for the kind words @CAM-Gerlach! There are more than a few similarities between good code and good documentation, I think 😄 I'll do my best to help out where and when I can.
Reacted by C.A.M. Gerlach- added 4 commits that reference this issue
on Dec 21, 2022 Thanks, looks like everything is merged!
Reacted by James Addison and Bisola OlasehindeReacted by James Addison- added a commit that references this issue
on Dec 28, 2022
Documentation
The are a couple of paragraphs about error-checking behaviour for
str.encodeandstr.decodethat appears in a few versions of Python and could potentially be improved in future versions.As found in the Python 3.9.14 documentation for
std.encode:The paragraph before that is fairly dense - there could be an opportunity to improve both of them.
Linked PRs