`zipfile` may write inappropriate extra fields to Local File Entries
Ninguém assumiu esta issue ainda.
- Linguagem predominante
- Python
- Estrelas
- 77.2k
- Forks
- 36k
- Métricas de merge de PRs
- Métricas de PR pendentes
Descrição
Bug report
Bug description:
Currently, ZipInfo.extra reads data from the Central Directory (CD), but writes to both the Local File Entry (LFE) and the Central Directory.
However, certain extra fields are layout-dependent or exclusive to the CD. Direct synchronization may cause corrupted or unreadable metadata.
1. Duplicated Zip64 Extra Fields (0x0001)
If ZipInfo.extra already contains a Zip64 extra field manually injected or copied from an existing archive, the LFE will contain duplicated Zip64 extra fields when written.
Since how to handle duplicated fields is undefined, ZIP tools may fail to read the correct zip64 data or treat the ZIP archive as corrupted due to inconsistency across LFE and CD.
For example:
import io
import struct
import zipfile
fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
zinfo = zipfile.ZipInfo('strfile')
zinfo.extra = (
b'\x01\x00\x10\x00'
b'\x00\x00\x00\x00\x00\x00\x00\x00'
b'\x00\x00\x00\x00\x00\x00\x00\x00'
)
with zh.open(zinfo, 'w', force_zip64=True) as zi:
pass
fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)
The result is double b'\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', while only one should exist.
2. Appended Zip64 Field Can Be Masked and Inaccessible
When ZipInfo.extra contains a misaligned or trailing field that ends abruptly, the appended Zip64 extra field is written immediately after it.
For example:
import io
import struct
import zipfile
fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
zinfo = zipfile.ZipInfo('strfile')
zinfo.extra = b'\x02\x00\x00\x01\x00\x00'
with zh.open(zinfo, 'w', force_zip64=True) as zi:
pass
fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)
The result is b'\x02\x00\x00\x01\x00\x00\x01\x00\x10\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00', which will be intepreted as having a field with ID \x0002 and length 256 (\x00\x01). The zip64 field will be treated as part of that field and is not readable.
It would be better to prepend (rather than append) the zip64 field, as how it is handled in central directory. Although this still gets invalid extra data, it will be easier to analysis and recover since the zip64 extra field is readable.
3. Redundant Info-ZIP Unicode Comment Fields (0x6375)
The Info-ZIP Unicode Comment field (0x6375) is strictly meant to map archive comments, which only exist in the CD.
Writing 0x6375 to the LFE is a violation of common layout expectations and bloats the local file header with dead data.
For example:
import io
import struct
import zipfile
fh = io.BytesIO()
with zipfile.ZipFile(fh, 'w') as zh:
zinfo = zipfile.ZipInfo('strfile')
zinfo.comment = b'1'
zinfo.extra = struct.pack(
'<HHBL5s', 0x6375, 11, 1, zipfile.crc32(b'1'), '1\u4e00'.encode('utf-8'))
with zh.open(zinfo, 'w') as zi:
pass
fh.seek(zinfo.header_offset)
entry = fh.read(zh.start_dir - zinfo.header_offset - zinfo.compress_size)
header = struct.unpack_from(zipfile.structFileHeader, entry)
extra = entry[-header[zipfile._FH_EXTRA_FIELD_LENGTH]:]
print(extra)
The result is the unicode comment field, while it should be empty.
Proposed Solution
Before writing to LFE, zipfile should strip or sanitize layout-incompatible fields from ZipInfo.extra.
If automatic fields (like Zip64) must be injected, they should be prepended rather than appended to prevent trailing parsing errors.
CPython versions tested on:
3.14, 3.16
Operating systems tested on:
No response
Linked PRs
- gh-153704
Guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Direção de pesquisa
Comece pelo caminho zipfile.ZipFile.open e pelo tratamento de ZipInfo.extra, incluindo force_zip64 e a construção do cabeçalho local em comparação com o cabeçalho central. Compare os exemplos relatados com as entradas de arquivo locais geradas; considera-se concluído quando campos incompatíveis com o layout não são escritos ali e os dados Zip64 adicionados automaticamente continuam analisáveis.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Stack de tecnologia
- python
- Domínio
- tooling
- Tipo de issue
- Bug
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Status de atividade
- Estagnada
- Clareza
- Razoavelmente clara
- Facilidade para iniciantes
- 25/100