python / python/cpython

Improve `encodings.normalize_encoding` behaviour or docs

Open
#136,702 9 comments 0 reactions 1 assignee View on GitHub

@StanFromIreland is already working on this.

Since Jul 16, 2025.

stdlib topic-unicode type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

Bug description:

This is to wait for https://github.com/python/cpython/issues/55531 / https://github.com/python/cpython/pull/136643

  • This function accepts bytes since https://github.com/python/cpython/commit/98297ee7815939b124156e438b22bd652d67b5db. This is however undocumented, and untested. I propose deprecating (& later removing) it, otherwise it should be documented and tested properly.

  • This function is documented as

    encoding should be ASCII only.

    However, this is only enforced for bytes (see point 1) with a ValueError, and for strings the characters are simply removed. We should be consistent and not depend on input type. I propose enforcing this for strings too with a ValueError.

cc @malemburg (Note, absolutely no rush for this one, we should get the performance issue/C implementation sorted out first :-)

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Linked PRs
  • gh-140030
  • gh-141345

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.