python / python/cpython

`codecs.encode` with `utf-*` encoding and errors returing `str` rejects surrogates blindly

Aperta
#127,305 7 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

interpreter-core stdlib type-feature
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Bug report

Bug description:

For codecs.encode,
with utf-* encoding, and a custom errors which returns str,
if you pass some characters that are not valid UTF characters (e.g. surrogates),
UnicodeEncodeError is just raised and there's not the expected (and documented) case
where the returned str is appended.

import codecs

ERRORS_NAME = "returning non-ascii"

# something being not encod-able via `utf-*`
BAD_UTF = "\uD800"  # the first high surrogate character


def register_repl_error(repl: str):
    def error_handle(exc: UnicodeEncodeError) -> tuple[str, int]:
        return (repl, exc.end)

    codecs.register_error(ERRORS_NAME, error_handle)


def encode_surrogate(encoding: str, repl: str):
    register_repl_error(repl)
    max_enc_len = 9
    pre = f"codecs.encode({BAD_UTF!r}, {encoding=:{max_enc_len}}) "
    try:
        res = codecs.encode(BAD_UTF, encoding, ERRORS_NAME)
    except UnicodeEncodeError as err:
        reason = err.reason
        print(pre + f"raises with {reason=}")

    else:
        print(pre + f"returns {res}")


NON_ASCII = "龍"  # loong in Chinese

## utf-*

for i in ('8', '16', '32', '16-le', '16-be'):
    encode_surrogate("utf-" + i, NON_ASCII)

print('-'*3)


# The following is some non-utf* encoding, which works fine

## cjk

### zh
for enc in ("gbk", "big5"):
    encode_surrogate(enc, NON_ASCII)

### jp
for enc in ("Shift_JIS", "EUC-JP"):
    encode_surrogate(enc, NON_ASCII)

Output:

codecs.encode('\ud800', encoding=utf-8    ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16   ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-32   ) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16-le) raises with reason='surrogates not allowed'
codecs.encode('\ud800', encoding=utf-16-be) raises with reason='surrogates not allowed'
---
codecs.encode('\ud800', encoding=gbk      ) returns b'\xfd\x88'
codecs.encode('\ud800', encoding=big5     ) returns b'\xc0s'
codecs.encode('\ud800', encoding=Shift_JIS) returns b'\x97\xb4'
codecs.encode('\ud800', encoding=EUC-JP   ) returns b'\xce\xb6'
CPython versions tested on:

3.9, 3.11, 3.12, 3.13, 3.14

Operating systems tested on:

Linux, Windows

Linked PRs
  • gh-139554

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia eseguendo il reproducer fornito per le utf-* encodings e confronta il suo comportamento con quello delle non-UTF encodings. Esamina il PR collegato gh-139554 e il percorso di gestione degli errori di codecs.encode; il lavoro è completato quando un handler personalizzato che restituisce uno str viene accettato per i surrogates non validi come documentato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.