python / python/cpython

UnicodeEncodeError during mime header parsing is unhandled in _header_value_parser.py

Ouverte
#132,794 7 commentaires 1 réaction 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

stdlib topic-email type-bug
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

My underlying issue is related to https://github.com/python/cpython/issues/79728#issuecomment-1093807522, where the parameter for a mime header has been encoded, but the bytes for a multi-byte character are split across multiple lines. python is unable to parse this.

Of course, ideally, the client would handle this properly, but I don't have any control over that.

In e91dee87edc, handling was added for UnicodeEncodeError, but UnicodeDecodeError is still unhandled and the entire parse raises at this point.

Here is a minimal reproduction of an email a client is sending. It would be nice if the filename itself could be correctly parsed, but at a minimum, I would like to be able to get the file contents:

import email.policy
import secrets
from datetime import datetime, timezone
import sys
from email import encoders, message_from_bytes, utils
from email.message import EmailMessage
from email.mime.application import MIMEApplication
from email.utils import make_msgid


def main() -> None:
    if sys.version_info[:2] < (3, 11):
        now = datetime.now(timezone.utc)
    else:
        from datetime import UTC

        now = datetime.now(UTC)

    from_ = "test@example.com"

    msg = EmailMessage(policy=email.policy.SMTP)
    msg["date"] = utils.format_datetime(now)

    msg["subject"] = "Mime param split bytes"

    msg["from"] = from_
    msg["message-id"] = make_msgid(domain=msg["from"].addresses[0].domain)

    msg["to"] = "other@example.com"

    msg.set_content("This is a test", "plain")

    if not msg.is_multipart():
        msg.make_mixed()

    attachment = MIMEApplication(
        secrets.token_bytes(256),
        "pdf",
        encoders.encode_base64,
        policy=email.policy.SMTP,
    )
    attachment.add_header("Content-Disposition", "attachment", filename="test")
    msg.attach(attachment)

    msg_bytes = msg.as_bytes()

    # split the bytes of a multi-byte character across lines.
    filename = "作業報告書【子】.pdf"
    filename_bytes = (
        ("%" + filename.encode("iso-2022-jp").hex("%").upper())
        .encode("ascii")
        .split(b"%52", 1)
    )
    msg_bytes = msg_bytes.replace(
        b'attachment; filename="test"',
        b"attachment;\r\n filename*0*=ISO-2022-JP''"
        + b"; \r\n filename*1*=".join(filename_bytes),
    )

    print(msg_bytes.decode())

    # trigger parsing the mime-part with the filename
    message_from_bytes(msg_bytes, policy=email.policy.SMTP).get_body("html")


if __name__ == "__main__":
    main()

I ran this with Python 3.8 - 3.14 all with the same result.

$ uv run --no-config --managed-python --python 3.13 run.py
Traceback (most recent call last):
  File "/tmp/run.py", line 65, in <module>
    main()
    ~~~~^^
  File "/tmp/run.py", line 62, in main
    message_from_bytes(msg_bytes, policy=email.policy.SMTP).get_body("html")
    ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/message.py", line 1054, in get_body
    for prio, part in self._find_body(self, preferencelist):
                      ~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/message.py", line 1025, in _find_body
    yield from self._find_body(subpart, preferencelist)
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/message.py", line 1014, in _find_body
    if part.is_attachment():
       ~~~~~~~~~~~~~~~~~~^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/message.py", line 1010, in is_attachment
    c_d = self.get('content-disposition')
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/message.py", line 507, in get
    return self.policy.header_fetch_parse(k, v)
           ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/policy.py", line 163, in header_fetch_parse
    return self.header_factory(name, value)
           ~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/headerregistry.py", line 604, in __call__
    return self[name](name, value)
           ~~~~~~~~~~^^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/headerregistry.py", line 192, in __new__
    cls.parse(value, kwds)
    ~~~~~~~~~^^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/headerregistry.py", line 449, in parse
    kwds['decoded'] = str(parse_tree)
                      ~~~^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/_header_value_parser.py", line 136, in __str__
    return ''.join(str(x) for x in self)
           ~~~~~~~^^^^^^^^^^^^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/_header_value_parser.py", line 136, in <genexpr>
    return ''.join(str(x) for x in self)
                   ~~~^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/_header_value_parser.py", line 814, in __str__
    for name, value in self.params:
                       ^^^^^^^^^^^
  File "/home/aclemons/.local/share/uv/python/cpython-3.13.3-linux-aarch64-gnu/lib/python3.13/email/_header_value_parser.py", line 799, in params
    value = value.decode(charset, 'surrogateescape')
UnicodeDecodeError: 'iso2022_jp' codec can't decode byte 0x3b in position 15: incomplete multibyte sequence
decoding with 'ISO-2022-JP' codec failed

If I patch my python and add UnicodeDecodeError to the except on line 800 in _header_value_parser.py (https://github.com/python/cpython/blob/e84624450dc0494271119018c699372245d724d9/Lib/email/_header_value_parser.py#L800), I can at least interact with the email, even if the attachment filename from the parameter is garbled.

I checked other tickets to see if this had already been reported. I mentioned the root problem here with the handling of parameters with split bytes (#79728), but I also saw #116705 which asked about handling of UnicodeDecodeError at exactly this point too.

Thank you.

CPython versions tested on:

3.13

Operating systems tested on:

Linux

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans Lib/email/_header_value_parser.py, autour de la propriété params et de l’échec de décodage signalé, puis exécutez la reproduction minimale avec message_from_bytes(...).get_body("html"). C’est terminé lorsque des paramètres MIME mal répartis ne provoquent plus de UnicodeDecodeError lors de l’analyse, tandis que le contenu de la pièce jointe reste accessible ; vérifiez le comportement obtenu du nom de fichier par rapport au rapport.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
backend
Type d'issue
Bug
Difficulté
3/5
Temps estimé
1-2 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
38/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.