Skip to content

gh-158633: say that subprocess text mode guesses the child's encoding - #158634

Open
AlKor13 wants to merge 1 commit into
python:mainfrom
AlKor13:docs-subprocess-text-encoding
Open

AlKor13 wants to merge 1 commit into
python:mainfrom
AlKor13:docs-subprocess-text-encoding

Conversation

@AlKor13

@AlKor13 AlKor13 commented Oct 3, 2026 •

Copy link
Copy Markdown

Documentation only. Closes #158633, and implements what @zooba described on #105312:

I would add a warning that it will guess the encoding of the child process and you'll get garbage if it guesses wrong, and you are always recommended to specify the encoding to use (on all platforms, though you are most likely to get into trouble on Windows).

One .. warning:: in frequently-used-arguments, after the paragraph on binary mode. Popen's section already refers to that one, so it is covered too.

Two things the current text leaves out, and the reason they are worth a warning rather than a note:

  • Where the failure appears. When the guess is wrong the output is either silently mojibake, or a UnicodeDecodeError raised while the stream is read — and that traceback comes from inside subprocess, not from the call that omitted encoding=, which is what sends people looking for a bug in the module.

  • Windows has more than one default at once. Measured on Windows 11, ANSI code page cp1251, console output code page cp866, so two children of one process need two different decoders:

    subprocess.run([sys.executable, "-c", "print('Тест')"], capture_output=True, text=True).stdout
    # 'Тест'   — the child wrote UTF-8
    
    subprocess.run("echo Тест", shell=True, text=True, stdout=subprocess.PIPE).stdout
    # '’Ґбв'       — the child wrote cp866

    Under UTF-8 mode, which :pep:686 makes the default in 3.15, the second becomes UnicodeDecodeError: 'utf-8' codec can't decode byte 0x92 rather than correct text — so the advice holds after 3.15, which is why the warning says so explicitly instead of presenting UTF-8 as the end of the problem.

The Windows paragraph keeps to @zooba's caveat that GetConsoleOutputCP is only right if you know the child uses it: it says a console program is read with that page, and that a Python child can instead be told what to write through PYTHONUTF8 / PYTHONIOENCODING in its env, which is knowledge rather than a guess.

No behaviour change and no new API, so I believe this is skip news; happy to add a blurb if you would rather have one. I have not signed the CLA yet — doing that now, and I will confirm here once it shows.

…coding

Text mode documents which encoding is used and not that the value is a guess
about the child process, so a wrong guess reads as a bug in this module: the
output is either silently mojibake, or a UnicodeDecodeError raised while the
stream is read, reported from inside subprocess rather than from the call that
is missing encoding=.

The warning keeps to what was asked for on pythongh-105312: the default is a guess,
pass encoding= on every platform, and on Windows more than one default is in
force at once -- the ANSI code page and the console output code page -- so a
console child is read with the console page while a Python child can be told
what to write through PYTHONUTF8 / PYTHONIOENCODING.
@python-cla-bot

python-cla-bot Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

All commit authors signed the Contributor License Agreement.

CLA signed

@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 cpython-previews | 🛠️ Build #34911658 | 📁 Comparing d1ea9b3 against main (1a85213)

  🔍 Preview build  

1 file changed
± library/subprocess.html

@AlKor13

AlKor13 commented Oct 3, 2026

Copy link
Copy Markdown
Author

CLA signed — the check is green now, and so are lint, Docs, Check EPUB, the HTML-ID check and the docs preview.

Rendered warning, for review without a local build: https://cpython-previews--158634.org.readthedocs.build/en/158634/library/subprocess.html#frequently-used-arguments — the five cross-references in it (UnicodeDecodeError, :pep:686, locale.getpreferredencoding, PYTHONUTF8, PYTHONIOENCODING) all resolve.

Happy to trim it if it reads long for one admonition; the Windows paragraph is the part I would cut first, since the first one carries the advice on its own.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review docs Documentation in the Doc dir skip news

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

Docs: subprocess text mode does not say that the default encoding is a guess about the child process

1 participant