python / python/cpython

`zipfile` may write inappropriate extra fields to Local File Entries

未关闭
#153,702 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

stdlib type-bug
主要语言
Python
星标
77.2k
派生
35.9k
PR 合并指标
PR 指标待抓取

描述

Bug report

Bug description:

Currently, ZipInfo.extra reads data from the Central Directory (CD), but writes to both the Local File Entry (LFE) and the Central Directory.

However, certain extra fields are layout-dependent or exclusive to the CD. Direct synchronization may cause corrupted or unreadable metadata.

1. Duplicated Zip64 Extra Fields (0x0001)

If ZipInfo.extra already contains a Zip64 extra field manually injected or copied from an existing archive, the LFE will contain duplicated Zip64 extra fields when written.

Since how to handle duplicated fields is undefined, ZIP tools may fail to read the correct zip64 data or treat the ZIP archive as corrupted due to inconsistency across LFE and CD.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.extra = (
        b'\x01\x00\x10\x00'
        b'\x00\x00\x00\x00\x00\x00\x00\x00'
        b'\x00\x00\x00\x00\x00\x00\x00\x00'
    )
    with zh.open(zinfo, 'w', force_zip64=True) as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is double b'\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', while only one should exist.

2. Appended Zip64 Field Can Be Masked and Inaccessible

When ZipInfo.extra contains a misaligned or trailing field that ends abruptly, the appended Zip64 extra field is written immediately after it.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.extra = b'\x02\x00\x00\x01\x00\x00'
    with zh.open(zinfo, 'w', force_zip64=True) as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is b'\x02\x00\x00\x01\x00\x00\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', which will be intepreted as having a field with ID \x0002 and length 256 (\x00\x01). The zip64 field will be treated as part of that field and is not readable.

It would be better to prepend (rather than append) the zip64 field, as how it is handled in central directory. Although this still gets invalid extra data, it will be easier to analysis and recover since the zip64 extra field is readable.

3. Redundant Info-ZIP Unicode Comment Fields (0x6375)

The Info-ZIP Unicode Comment field (0x6375) is strictly meant to map archive comments, which only exist in the CD.

Writing 0x6375 to the LFE is a violation of common layout expectations and bloats the local file header with dead data.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.comment = b'1'
    zinfo.extra = struct.pack(
        '<HHBL5s', 0x6375, 11, 1, zipfile.crc32(b'1'), '1\u4e00'.encode('utf-8'))
    with zh.open(zinfo, 'w') as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is the unicode comment field, while it should be empty.

Proposed Solution

Before writing to LFE, zipfile should strip or sanitize layout-incompatible fields from ZipInfo.extra.

If automatic fields (like Zip64) must be injected, they should be prepended rather than appended to prevent trailing parsing errors.

CPython versions tested on:

3.14, 3.16

Operating systems tested on:

No response

Linked PRs
  • gh-153704

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从 zipfile.ZipFile.open 路径以及 ZipInfo.extra 的处理开始,包括 force_zip64 和本地标头与中央标头的构造。将报告的示例与生成的本地文件条目进行比较;完成的标准是,不兼容布局的字段不会写入其中,并且自动添加的 Zip64 数据仍可解析。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
tooling
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。