tarfile: member reads can propagate underlying compression exceptions instead of wrapping them
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
Bug report
Bug description:
Summary
TL;DR: tarfile can sometimes propagate an underlying compression error with its raw exception, rather than wrapping that exception in a tarfile.ReadError or similar. This makes exception handling for tarfiles a little bit unwieldy, since consumers have to remember to catch both tarfile exceptions and any underlying compression library exceptions.
In most cases, this doesn't happen, since tarfile catches and wraps decompression errors during the initial (header) read for each member. However, if the underlying compressed member payload is malformed after the header read, those subsequent decompression errors aren't caught. This doesn't happen in normal operation, but can happen if the user's encoder is buggy or they've somehow manipulated the compressed stream past the tar header.
I believe there's no security risk to this, it's just a (minor) nuisance to downstream users of the tarfile API who need to catch additional exception types. https://github.com/python/cpython/issues/83220 already handled a variant of this 🙂
The following is an (AI assisted) MRE, showing how four different underlying compression libraries have their exceptions leak through:
#!/usr/bin/env python3
from __future__ import annotations
import bz2
from collections.abc import Callable
from dataclasses import dataclass
import gzip
import io
import lzma
import sys
import tarfile
import zlib
from compression import zstd
MEMBER_NAME = "payload.bin"
PAYLOAD = b"A" * 1024
READ_CHUNK_SIZE = 16
class ChunkedBytesIO(io.BytesIO):
"""Limit compressed reads so corruption is reached after open()."""
def read(self, size: int = -1) -> bytes:
if size < 0 or size > READ_CHUNK_SIZE:
size = READ_CHUNK_SIZE
return super().read(size)
@dataclass(frozen=True)
class Case:
label: str
mode: str
compress: Callable[[bytes], bytes]
signature: bytes
corrupt_offset: int
expected_error: type[BaseException]
expected_name: str
replacement: int | None = None
def gzip_compress(data: bytes) -> bytes:
return gzip.compress(data, mtime=0)
CASES = (
# Byte 10 starts the DEFLATE stream. BTYPE=3 is reserved, so 0b111 is
# an invalid first block header.
Case(
"gzip",
"r:gz",
gzip_compress,
b"\x1f\x8b\x08",
10,
zlib.error,
"zlib.error",
replacement=0b111,
),
# Flip a byte in the first bzip2 block header/data.
Case(
"bzip2",
"r:bz2",
bz2.compress,
b"BZh",
16,
OSError,
"OSError",
),
# Flip a byte in the first XZ block header.
Case(
"xz/lzma",
"r:xz",
lzma.compress,
b"\xfd7zXZ\x00",
13,
lzma.LZMAError,
"lzma.LZMAError",
),
# Flip the Zstandard frame header descriptor.
Case(
"zstandard",
"r:zst",
zstd.compress,
b"\x28\xb5\x2f\xfd",
4,
zstd.ZstdError,
"compression.zstd.ZstdError",
),
)
def make_tar() -> bytes:
buffer = io.BytesIO()
with tarfile.open(fileobj=buffer, mode="w:") as archive:
member = tarfile.TarInfo(MEMBER_NAME)
member.size = len(PAYLOAD)
archive.addfile(member, io.BytesIO(PAYLOAD))
return buffer.getvalue()
def make_corrupt_archive(tar_bytes: bytes, case: Case) -> bytes:
# Concatenated compressed streams decode as one continuous byte stream.
# The first contains the tar header and half the payload. Corruption is in
# the second, so opening succeeds but reading the full payload fails.
split_at = 512 + len(PAYLOAD) // 2
first_stream = case.compress(tar_bytes[:split_at])
second_stream = bytearray(case.compress(tar_bytes[split_at:]))
assert second_stream.startswith(case.signature)
if case.replacement is None:
second_stream[case.corrupt_offset] ^= 0xFF
else:
second_stream[case.corrupt_offset] = case.replacement
return first_stream + second_stream
def reproduce(tar_bytes: bytes, case: Case) -> bool:
corrupt_archive = make_corrupt_archive(tar_bytes, case)
with tarfile.open(
fileobj=ChunkedBytesIO(corrupt_archive), mode=case.mode
) as archive:
# next() returns the first member cached during open(), without scanning
# later headers and encountering the corruption early.
member = archive.next()
assert member is not None and member.name == MEMBER_NAME
extracted = archive.extractfile(member)
assert extracted is not None
try:
with extracted:
extracted.read()
except case.expected_error as error:
actual_name = f"{type(error).__module__}.{type(error).__name__}"
print(f"{case.label}: leaked {actual_name}")
print(f" message: {error}")
print(f" expected tarfile.ReadError, not {case.expected_name}")
return True
except Exception as error:
actual_name = f"{type(error).__module__}.{type(error).__name__}"
print(f"{case.label}: got unexpected {actual_name}: {error}")
return False
print(f"{case.label}: corrupt member unexpectedly read successfully")
return False
def main() -> int:
print(sys.version)
tar_bytes = make_tar()
reproduced = [reproduce(tar_bytes, case) for case in CASES]
print(f"\nreproduced {sum(reproduced)}/{len(CASES)} exception leaks")
return 0 if all(reproduced) else 1
if __name__ == "__main__":
raise SystemExit(main())
Other context
See https://github.com/pypi/warehouse/pull/20415 for context.
Related: https://github.com/python/cpython/issues/83220
Note: this issue is 100% human written, but the MRE script was generated by Codex.
CPython versions tested on:
CPython main branch
Operating systems tested on:
No response
Linked PRs
- gh-156143
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从提供的复现程序开始,跟踪 tarfile.open()、archive.extractfile(member) 和 extracted.read() 的路径。完成的标准是:成员读取期间的损坏由 tarfile 级别的 ReadError 表示,而不是底层的压缩异常,并覆盖 gzip、bzip2、xz/lzma 和 zstandard 情况。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- api
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 25/100