python / python/cpython

tarfile: member reads can propagate underlying compression exceptions instead of wrapping them

Đang mở
#156,057 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

stdlib type-bug
Ngôn ngữ chính
Python
Star
77.2k
Fork
35.9k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

Bug report

Bug description:

Summary

TL;DR: tarfile can sometimes propagate an underlying compression error with its raw exception, rather than wrapping that exception in a tarfile.ReadError or similar. This makes exception handling for tarfiles a little bit unwieldy, since consumers have to remember to catch both tarfile exceptions and any underlying compression library exceptions.

In most cases, this doesn't happen, since tarfile catches and wraps decompression errors during the initial (header) read for each member. However, if the underlying compressed member payload is malformed after the header read, those subsequent decompression errors aren't caught. This doesn't happen in normal operation, but can happen if the user's encoder is buggy or they've somehow manipulated the compressed stream past the tar header.

I believe there's no security risk to this, it's just a (minor) nuisance to downstream users of the tarfile API who need to catch additional exception types. https://github.com/python/cpython/issues/83220 already handled a variant of this 🙂

The following is an (AI assisted) MRE, showing how four different underlying compression libraries have their exceptions leak through:

#!/usr/bin/env python3

from __future__ import annotations

import bz2
from collections.abc import Callable
from dataclasses import dataclass
import gzip
import io
import lzma
import sys
import tarfile
import zlib

from compression import zstd


MEMBER_NAME = "payload.bin"
PAYLOAD = b"A" * 1024
READ_CHUNK_SIZE = 16


class ChunkedBytesIO(io.BytesIO):
    """Limit compressed reads so corruption is reached after open()."""

    def read(self, size: int = -1) -> bytes:
        if size < 0 or size > READ_CHUNK_SIZE:
            size = READ_CHUNK_SIZE
        return super().read(size)


@dataclass(frozen=True)
class Case:
    label: str
    mode: str
    compress: Callable[[bytes], bytes]
    signature: bytes
    corrupt_offset: int
    expected_error: type[BaseException]
    expected_name: str
    replacement: int | None = None


def gzip_compress(data: bytes) -> bytes:
    return gzip.compress(data, mtime=0)


CASES = (
    # Byte 10 starts the DEFLATE stream. BTYPE=3 is reserved, so 0b111 is
    # an invalid first block header.
    Case(
        "gzip",
        "r:gz",
        gzip_compress,
        b"\x1f\x8b\x08",
        10,
        zlib.error,
        "zlib.error",
        replacement=0b111,
    ),
    # Flip a byte in the first bzip2 block header/data.
    Case(
        "bzip2",
        "r:bz2",
        bz2.compress,
        b"BZh",
        16,
        OSError,
        "OSError",
    ),
    # Flip a byte in the first XZ block header.
    Case(
        "xz/lzma",
        "r:xz",
        lzma.compress,
        b"\xfd7zXZ\x00",
        13,
        lzma.LZMAError,
        "lzma.LZMAError",
    ),
    # Flip the Zstandard frame header descriptor.
    Case(
        "zstandard",
        "r:zst",
        zstd.compress,
        b"\x28\xb5\x2f\xfd",
        4,
        zstd.ZstdError,
        "compression.zstd.ZstdError",
    ),
)


def make_tar() -> bytes:
    buffer = io.BytesIO()
    with tarfile.open(fileobj=buffer, mode="w:") as archive:
        member = tarfile.TarInfo(MEMBER_NAME)
        member.size = len(PAYLOAD)
        archive.addfile(member, io.BytesIO(PAYLOAD))
    return buffer.getvalue()


def make_corrupt_archive(tar_bytes: bytes, case: Case) -> bytes:
    # Concatenated compressed streams decode as one continuous byte stream.
    # The first contains the tar header and half the payload. Corruption is in
    # the second, so opening succeeds but reading the full payload fails.
    split_at = 512 + len(PAYLOAD) // 2
    first_stream = case.compress(tar_bytes[:split_at])
    second_stream = bytearray(case.compress(tar_bytes[split_at:]))

    assert second_stream.startswith(case.signature)
    if case.replacement is None:
        second_stream[case.corrupt_offset] ^= 0xFF
    else:
        second_stream[case.corrupt_offset] = case.replacement

    return first_stream + second_stream


def reproduce(tar_bytes: bytes, case: Case) -> bool:
    corrupt_archive = make_corrupt_archive(tar_bytes, case)

    with tarfile.open(
        fileobj=ChunkedBytesIO(corrupt_archive), mode=case.mode
    ) as archive:
        # next() returns the first member cached during open(), without scanning
        # later headers and encountering the corruption early.
        member = archive.next()
        assert member is not None and member.name == MEMBER_NAME

        extracted = archive.extractfile(member)
        assert extracted is not None

        try:
            with extracted:
                extracted.read()
        except case.expected_error as error:
            actual_name = f"{type(error).__module__}.{type(error).__name__}"
            print(f"{case.label}: leaked {actual_name}")
            print(f"  message: {error}")
            print(f"  expected tarfile.ReadError, not {case.expected_name}")
            return True
        except Exception as error:
            actual_name = f"{type(error).__module__}.{type(error).__name__}"
            print(f"{case.label}: got unexpected {actual_name}: {error}")
            return False

    print(f"{case.label}: corrupt member unexpectedly read successfully")
    return False


def main() -> int:
    print(sys.version)
    tar_bytes = make_tar()
    reproduced = [reproduce(tar_bytes, case) for case in CASES]
    print(f"\nreproduced {sum(reproduced)}/{len(CASES)} exception leaks")
    return 0 if all(reproduced) else 1


if __name__ == "__main__":
    raise SystemExit(main())

Other context

See https://github.com/pypi/warehouse/pull/20415 for context.

Related: https://github.com/python/cpython/issues/83220

Note: this issue is 100% human written, but the MRE script was generated by Codex.

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Linked PRs
  • gh-156143

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu với reproducer được cung cấp và lần theo đường đi của tarfile.open(), archive.extractfile(member) và extracted.read(). Công việc được xem là hoàn tất khi lỗi hỏng dữ liệu trong quá trình đọc member được biểu diễn bằng ReadError ở cấp tarfile thay vì ngoại lệ nén bên dưới, với độ bao phủ cho các trường hợp gzip, bzip2, xz/lzma và zstandard.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
api
Loại issue
Lỗi
Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.