Improve `encodings.normalize_encoding` behaviour or docs
@StanFromIreland is already working on this.
Since Jul 16, 2025.
- Dominant language
- Python
- Stars
- 77.2k
- Forks
- 35.9k
- PR merge metrics
- PR metrics pending
Description
Bug report
Bug description:
This is to wait for https://github.com/python/cpython/issues/55531 / https://github.com/python/cpython/pull/136643
-
This function accepts bytes since https://github.com/python/cpython/commit/98297ee7815939b124156e438b22bd652d67b5db. This is however undocumented, and untested. I propose deprecating (& later removing) it, otherwise it should be documented and tested properly.
-
This function is documented as
encoding should be ASCII only.
However, this is only enforced for bytes (see point 1) with a
ValueError, and for strings the characters are simply removed. We should be consistent and not depend on input type. I propose enforcing this for strings too with aValueError.
cc @malemburg (Note, absolutely no rush for this one, we should get the performance issue/C implementation sorted out first :-)
CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs
- gh-140030
- gh-141345
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.