lzma: unbounded dictionary-size allocation is a memory DoS (zipfile, tarfile, lzma.decompress)
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 35.9k
- PR 合併指標
- PR 指標待擷取
描述
Bug: LZMA decompression accepts an attacker-controlled dictionary size with no bound, and liblzma allocates the full dictionary up front when a decompressor is created. A few bytes of input can therefore force a multi-gigabyte allocation (memory denial of service).
Attack surfaces (both reported by OSS-Fuzz, both public):
zipfilewith an LZMA-compressed member: the 5-byte LZMA properties header in the member data declares the dictionary. A 1088-byte zip whose member declares a ~2 GiB dictionary OOMs the process. https://issues.oss-fuzz.com/issues/495472861tarfile's compression-detection loop: a 13-byte input whose first byte is a valid LZMA-alone properties byte (0x66) and whose remaining bytes encode a ~4 GiB dictionary is tried as LZMA; liblzma allocates 4 GiB before failing. https://issues.oss-fuzz.com/issues/482161128
Also reachable directly: lzma.decompress() / LZMADecompressor(format=FORMAT_AUTO) on untrusted bytes, and LZMADecompressor(FORMAT_RAW, filters=[{"dict_size": ...}]).
Root cause: two gaps:
parse_filter_spec_lzma()accepts anydict_size(up to UINT32_MAX) for raw LZMA1/LZMA2 filter chains; liblzma allocates that dictionary at decompressor creation.LZMADecompressordefaultsmemlimittoUINT64_MAX, so the FORMAT_ALONE/FORMAT_AUTO paths (where the header is parsed inside liblzma) are also unbounded.
Proposed fix (PR to follow):
- Reject
dict_size > LZMA_DICT_SIZE_MAX(1.5 GiB, liblzma's own maximum) inparse_filter_spec_lzma()withLZMAError. - Default the decompressor
memlimittoLZMA_DICT_SIZE_MAXinstead ofUINT64_MAX(callers may still pass an explicitmemlimit).
No legitimate streams are affected: liblzma presets cap at 64 MiB dictionaries, and 1.5 GiB is liblzma's own maximum supported dictionary size.
Reported by OSS-Fuzz (ClusterFuzz). Reproducers attached to the issues above.
Linked PRs
- gh-155650
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 parse_filter_spec_lzma() 和 LZMADecompressor 開始,然後檢查 issue 中描述的 zipfile 和 tarfile 路徑。使用連結的 OSS-Fuzz 重現器進行驗證,確保過大的字典不再導致無界配置,同時正常的 LZMA 串流和明確的 memlimit 行為維持不變。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- security
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 停滯
- 描述清晰度
- 描述清楚
- 新手友好度
- 35/100