python / python/cpython

mailbox.MH.get_sequences() unnecessarily materializes large sequence ranges

Open
#156,379 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

performance stdlib topic-email type-feature
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Feature or enhancement

Proposal:

Hello Python Security Response Team,

I am reporting a resource-exhaustion issue in CPython's mailbox.MH implementation. MH.get_sequences() materializes every integer in a range from the .mh_sequences metadata file before filtering the result against the message keys that actually exist. A normal MH.get_message() call invokes get_sequences(), so a tiny metadata file can cause a large allocation while loading an otherwise ordinary message.
Reproduction:

Please see the attached plain-text proof-of-concept, cpython-mh-sequence-repro.py. It creates a temporary MH mailbox containing one message and a small .mh_sequences file, then calls MH.get_message(1).

On macOS with Python 3.14.6:

$ python3 cpython-mh-sequence-repro.py --stop 5000000
python=3.14.6
stop=5000000 metadata_bytes=18
elapsed=0.259s maxrss=451297280 outcome=message_loaded

Negative control:

$ python3 cpython-mh-sequence-repro.py --stop 5
python=3.14.6
stop=5 metadata_bytes=12
elapsed=0.001s maxrss=23674880 outcome=message_loaded

The latest main ASAN build also reproduced the behavior with a one-million range: 18 bytes of metadata reached approximately 223 MB maximum RSS while loading the single message. No ASAN diagnostic occurred; this report is not claiming a memory-safety issue.

Root cause:

In Lib/mailbox.py, MH.get_sequences() effectively executes keys.update(range(start, stop + 1)) and only afterwards filters the result against all_keys. The final result only contains existing mailbox keys, so the declared interval does not need to be enumerated.

Suggested mitigation:

Intersect each declared range with all_keys instead of enumerating the interval, for example:

if start <= stop:
    keys.update(key for key in all_keys if start <= key <= stop)

The fix should preserve existing malformed-input behavior, ordering, and empty-sequence removal, and should add a regression test using a very large range with a small mailbox.

I searched the CPython issue and pull-request tracker for get_sequences, .mh_sequences, keys.update(range, and MH mailbox DoS. I did not find a matching report. Existing MH issues concerning Claws Mail sequence-file formats, missing .mh_sequences files, and replacement data loss are different behaviors.

POC is here.

#!/usr/bin/env python3
"""Measure MH sequence-range materialization from a tiny mailbox file."""

from __future__ import annotations

import argparse
from pathlib import Path
import mailbox
import resource
import sys
import tempfile
import time


def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--stop", type=int, default=1_000_000)
    args = parser.parse_args()
    with tempfile.TemporaryDirectory(prefix="cpython-mh-") as tmp:
        root = Path(tmp)
        (root / "1").write_bytes(b"Subject: canary\n\nbody\n")
        (root / ".mh_sequences").write_text(
            f"unseen: 1-{args.stop}\n", encoding="ascii"
        )
        box = mailbox.MH(root)
        start = time.monotonic()
        try:
            box.get_message(1)
            outcome = "message_loaded"
        except BaseException as exc:
            outcome = f"{type(exc).__name__}: {exc}"
        elapsed = time.monotonic() - start
        maxrss = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss
        print(f"python={sys.version.split()[0]}")
        print(f"stop={args.stop} metadata_bytes={(root / '.mh_sequences').stat().st_size}")
        print(f"elapsed={elapsed:.3f}s maxrss={maxrss} outcome={outcome}")
        return 1 if args.stop >= 1_000_000 else 0


if __name__ == "__main__":
    raise SystemExit(main())
Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs
  • gh-156442

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The affected implementation is Lib/mailbox.py, specifically MH.get_sequences(); start there and inspect the existing mailbox tests. Add a regression test using a very large range with a small mailbox, preserving malformed-input behavior, ordering, and empty-sequence removal.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, security
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.