python / python/cpython

`codecs.encode` with `utf-*` encoding and errors returing `str` rejects surrogates blindly

Ouverte
#127,305 7 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

interpreter-core stdlib type-feature
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

For codecs.encode,
with utf-* encoding, and a custom errors which returns str,
if you pass some characters that are not valid UTF characters (e.g. surrogates),
UnicodeEncodeError is just raised and there's not the expected (and documented) case
where the returned str is appended.

import codecs

ERRORS_NAME = "returning non-ascii"

# something being not encod-able via `utf-*`
BAD_UTF = "\uD800"  # the first high surrogate character


def register_repl_error(repl: str):
    def error_handle(exc: UnicodeEncodeError) -> tuple[str, int]:
        return (repl, exc.end)

    codecs.register_error(ERRORS_NAME, error_handle)


def encode_surrogate(encoding: str, repl: str):
    register_repl_error(repl)
    max_enc_len = 9
    pre = f"codecs.encode({BAD_UTF!r}, {encoding=:{max_enc_len}}) "
    try:
        res = codecs.encode(BAD_UTF, encoding, ERRORS_NAME)
    except UnicodeEncodeError as err:
        reason = err.reason
        print(pre + f"raises with {reason=}")

    else:
        print(pre + f"returns {res}")


NON_ASCII = "龍"  # loong in Chinese

## utf-*

for i in ('8', '16', '32', '16-le', '16-be'):
    encode_surrogate("utf-" + i, NON_ASCII)

print('-'*3)


# The following is some non-utf* encoding, which works fine

## cjk

### zh
for enc in ("gbk", "big5"):
    encode_surrogate(enc, NON_ASCII)

### jp
for enc in ("Shift_JIS", "EUC-JP"):
    encode_surrogate(enc, NON_ASCII)

Output:

codecs.encode('\ud800', encoding=utf-8    ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16   ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-32   ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16-le) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16-be) raises with reason='surrogates not allowed'
---
codecs.encode('\ud800', encoding=gbk      ) returns b'\xfd\x88'
codecs.encode('\ud800', encoding=big5     ) returns b'\xc0s'
codecs.encode('\ud800', encoding=Shift_JIS) returns b'\x97\xb4'
codecs.encode('\ud800', encoding=EUC-JP   ) returns b'\xce\xb6'
CPython versions tested on:

3.9, 3.11, 3.12, 3.13, 3.14

Operating systems tested on:

Linux, Windows

Linked PRs
  • gh-139554

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par exécuter le reproducer fourni pour les utf-* encodings et comparez son comportement avec celui des non-UTF encodings. Examinez le PR lié gh-139554 et le chemin de gestion des erreurs de codecs.encode ; c’est terminé lorsqu’un handler personnalisé renvoyant un str est accepté pour les surrogates invalides, comme documenté.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
backend
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
25/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.