python / python/cpython

`zipfile` may write inappropriate extra fields to Local File Entries

Open
#153,702 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

stdlib type-bug
Dominant language
Python
Stars
77.2k
Forks
35.9k
PR merge metrics
PR metrics pending

Description

Bug report

Bug description:

Currently, ZipInfo.extra reads data from the Central Directory (CD), but writes to both the Local File Entry (LFE) and the Central Directory.

However, certain extra fields are layout-dependent or exclusive to the CD. Direct synchronization may cause corrupted or unreadable metadata.

1. Duplicated Zip64 Extra Fields (0x0001)

If ZipInfo.extra already contains a Zip64 extra field manually injected or copied from an existing archive, the LFE will contain duplicated Zip64 extra fields when written.

Since how to handle duplicated fields is undefined, ZIP tools may fail to read the correct zip64 data or treat the ZIP archive as corrupted due to inconsistency across LFE and CD.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.extra = (
        b'\x01\x00\x10\x00'
        b'\x00\x00\x00\x00\x00\x00\x00\x00'
        b'\x00\x00\x00\x00\x00\x00\x00\x00'
    )
    with zh.open(zinfo, 'w', force_zip64=True) as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is double b'\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', while only one should exist.

2. Appended Zip64 Field Can Be Masked and Inaccessible

When ZipInfo.extra contains a misaligned or trailing field that ends abruptly, the appended Zip64 extra field is written immediately after it.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.extra = b'\x02\x00\x00\x01\x00\x00'
    with zh.open(zinfo, 'w', force_zip64=True) as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is b'\x02\x00\x00\x01\x00\x00\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', which will be intepreted as having a field with ID \x0002 and length 256 (\x00\x01). The zip64 field will be treated as part of that field and is not readable.

It would be better to prepend (rather than append) the zip64 field, as how it is handled in central directory. Although this still gets invalid extra data, it will be easier to analysis and recover since the zip64 extra field is readable.

3. Redundant Info-ZIP Unicode Comment Fields (0x6375)

The Info-ZIP Unicode Comment field (0x6375) is strictly meant to map archive comments, which only exist in the CD.

Writing 0x6375 to the LFE is a violation of common layout expectations and bloats the local file header with dead data.

For example:

import io
import struct
import zipfile

fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
    zinfo = zipfile.ZipInfo('strfile')
    zinfo.comment = b'1'
    zinfo.extra = struct.pack(
        '<HHBL5s', 0x6375, 11, 1, zipfile.crc32(b'1'), '1\u4e00'.encode('utf-8'))
    with zh.open(zinfo, 'w') as zi:
        pass

fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)

The result is the unicode comment field, while it should be empty.

Proposed Solution

Before writing to LFE, zipfile should strip or sanitize layout-incompatible fields from ZipInfo.extra.

If automatic fields (like Zip64) must be injected, they should be prepended rather than appended to prevent trailing parsing errors.

CPython versions tested on:

3.14, 3.16

Operating systems tested on:

No response

Linked PRs
  • gh-153704

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the zipfile.ZipFile.open path and the handling of ZipInfo.extra, including force_zip64 and the local-versus-central header construction. Compare the reported examples with the generated local file entries; done means layout-incompatible fields are not written there and automatically added Zip64 data remains parseable.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
tooling
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.