python / python/cpython

mailbox.MH.get_sequences() unnecessarily materializes large sequence ranges

Aperta
#156,379 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

performance stdlib topic-email type-feature
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Feature or enhancement

Proposal:

Hello Python Security Response Team,

I am reporting a resource-exhaustion issue in CPython's mailbox.MH implementation. MH.get_sequences() materializes every integer in a range from the .mh_sequences metadata file before filtering the result against the message keys that actually exist. A normal MH.get_message() call invokes get_sequences(), so a tiny metadata file can cause a large allocation while loading an otherwise ordinary message.
Reproduction:

Please see the attached plain-text proof-of-concept, cpython-mh-sequence-repro.py. It creates a temporary MH mailbox containing one message and a small .mh_sequences file, then calls MH.get_message(1).

On macOS with Python 3.14.6:

$ python3 cpython-mh-sequence-repro.py --stop 5000000
python=3.14.6
stop=5000000 metadata_bytes=18
elapsed=0.259s maxrss=451297280 outcome=message_loaded

Negative control:

$ python3 cpython-mh-sequence-repro.py --stop 5
python=3.14.6
stop=5 metadata_bytes=12
elapsed=0.001s maxrss=23674880 outcome=message_loaded

The latest main ASAN build also reproduced the behavior with a one-million range: 18 bytes of metadata reached approximately 223 MB maximum RSS while loading the single message. No ASAN diagnostic occurred; this report is not claiming a memory-safety issue.

Root cause:

In Lib/mailbox.py, MH.get_sequences() effectively executes keys.update(range(start, stop + 1)) and only afterwards filters the result against all_keys. The final result only contains existing mailbox keys, so the declared interval does not need to be enumerated.

Suggested mitigation:

Intersect each declared range with all_keys instead of enumerating the interval, for example:

if start <= stop:
    keys.update(key for key in all_keys if start <= key <= stop)

The fix should preserve existing malformed-input behavior, ordering, and empty-sequence removal, and should add a regression test using a very large range with a small mailbox.

I searched the CPython issue and pull-request tracker for get_sequences, .mh_sequences, keys.update(range, and MH mailbox DoS. I did not find a matching report. Existing MH issues concerning Claws Mail sequence-file formats, missing .mh_sequences files, and replacement data loss are different behaviors.

POC is here.

#!/usr/bin/env python3
"""Measure MH sequence-range materialization from a tiny mailbox file."""

from __future__ import annotations

import argparse
from pathlib import Path
import mailbox
import resource
import sys
import tempfile
import time


def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--stop", type=int, default=1_000_000)
    args = parser.parse_args()
    with tempfile.TemporaryDirectory(prefix="cpython-mh-") as tmp:
        root = Path(tmp)
        (root / "1").write_bytes(b"Subject: canary\n\nbody\n")
        (root / ".mh_sequences").write_text(
            f"unseen: 1-{args.stop}\n", encoding="ascii"
        )
        box = mailbox.MH(root)
        start = time.monotonic()
        try:
            box.get_message(1)
            outcome = "message_loaded"
        except BaseException as exc:
            outcome = f"{type(exc).__name__}: {exc}"
        elapsed = time.monotonic() - start
        maxrss = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
        print(f"python={sys.version.split()[0]}")
        print(f"stop={args.stop} metadata_bytes={(root / '.mh_sequences').stat().st_size}")
        print(f"elapsed={elapsed:.3f}s maxrss={maxrss} outcome={outcome}")
        return 1 if args.stop >= 1_000_000 else 0


if __name__ == "__main__":
    raise SystemExit(main())
Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-156442

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

L'implementazione interessata si trova in Lib/mailbox.py, nello specifico in MH.get_sequences(); inizia da lì e ispeziona i test esistenti di mailbox. Aggiungi un test di regressione che utilizzi un intervallo molto grande con una mailbox piccola, preservando il comportamento per gli input malformati, l'ordinamento e la rimozione delle sequenze vuote.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
backend, security
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Ferma
Chiarezza
Specificata chiaramente
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.