python / python/cpython

mailbox.MH.get_sequences() unnecessarily materializes large sequence ranges

Offen
#156,379 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

performance stdlib topic-email type-feature
Vorherrschende Sprache
Python
Sterne
77.2k
Forks
35.9k
PR-Merge-Kennzahlen
PR-Kennzahlen ausstehend

Beschreibung

Feature or enhancement

Proposal:

Hello Python Security Response Team,

I am reporting a resource-exhaustion issue in CPython's mailbox.MH implementation. MH.get_sequences() materializes every integer in a range from the .mh_sequences metadata file before filtering the result against the message keys that actually exist. A normal MH.get_message() call invokes get_sequences(), so a tiny metadata file can cause a large allocation while loading an otherwise ordinary message.
Reproduction:

Please see the attached plain-text proof-of-concept, cpython-mh-sequence-repro.py. It creates a temporary MH mailbox containing one message and a small .mh_sequences file, then calls MH.get_message(1).

On macOS with Python 3.14.6:

$ python3 cpython-mh-sequence-repro.py --stop 5000000
python=3.14.6
stop=5000000 metadata_bytes=18
elapsed=0.259s maxrss=451297280 outcome=message_loaded

Negative control:

$ python3 cpython-mh-sequence-repro.py --stop 5
python=3.14.6
stop=5 metadata_bytes=12
elapsed=0.001s maxrss=23674880 outcome=message_loaded

The latest main ASAN build also reproduced the behavior with a one-million range: 18 bytes of metadata reached approximately 223 MB maximum RSS while loading the single message. No ASAN diagnostic occurred; this report is not claiming a memory-safety issue.

Root cause:

In Lib/mailbox.py, MH.get_sequences() effectively executes keys.update(range(start, stop + 1)) and only afterwards filters the result against all_keys. The final result only contains existing mailbox keys, so the declared interval does not need to be enumerated.

Suggested mitigation:

Intersect each declared range with all_keys instead of enumerating the interval, for example:

if start <= stop:
    keys.update(key for key in all_keys if start <= key <= stop)

The fix should preserve existing malformed-input behavior, ordering, and empty-sequence removal, and should add a regression test using a very large range with a small mailbox.

I searched the CPython issue and pull-request tracker for get_sequences, .mh_sequences, keys.update(range, and MH mailbox DoS. I did not find a matching report. Existing MH issues concerning Claws Mail sequence-file formats, missing .mh_sequences files, and replacement data loss are different behaviors.

POC is here.

#!/usr/bin/env python3
"""Measure MH sequence-range materialization from a tiny mailbox file."""

from __future__ import annotations

import argparse
from pathlib import Path
import mailbox
import resource
import sys
import tempfile
import time


def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--stop", type=int, default=1_000_000)
    args = parser.parse_args()
    with tempfile.TemporaryDirectory(prefix="cpython-mh-") as tmp:
        root = Path(tmp)
        (root / "1").write_bytes(b"Subject: canary\n\nbody\n")
        (root / ".mh_sequences").write_text(
            f"unseen: 1-{args.stop}\n", encoding="ascii"
        )
        box = mailbox.MH(root)
        start = time.monotonic()
        try:
            box.get_message(1)
            outcome = "message_loaded"
        except BaseException as exc:
            outcome = f"{type(exc).__name__}: {exc}"
        elapsed = time.monotonic() - start
        maxrss = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
        print(f"python={sys.version.split()[0]}")
        print(f"stop={args.stop} metadata_bytes={(root / '.mh_sequences').stat().st_size}")
        print(f"elapsed={elapsed:.3f}s maxrss={maxrss} outcome={outcome}")
        return 1 if args.stop >= 1_000_000 else 0


if __name__ == "__main__":
    raise SystemExit(main())
Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-156442

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Die betroffene Implementierung befindet sich in Lib/mailbox.py, insbesondere in MH.get_sequences(); beginne dort und prüfe die bestehenden Mailbox-Tests. Füge einen Regressionstest hinzu, der einen sehr großen Bereich mit einer kleinen Mailbox verwendet und dabei das Verhalten bei ungültigen Eingaben, die Reihenfolge und das Entfernen leerer Sequenzen beibehält.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
backend, security
Issue-Typ
Bug
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Veraltet
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.