python / python/cpython

Enhance `zipfile` to support decoding comments using `metadata_encoding`

Aperta
#152,929 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

stdlib type-feature
Lingua principale
Python
Stelle
77.2k
Fork
35.9k
Metriche di merge delle PR
Metriche PR in attesa

Descrizione

Feature or enhancement

Proposal:

In the current implementation, ZipFile accepts a metadata_encoding parameter to decode member filenames automatically into strings. However, {ZipFile,ZipInfo}.comment remains as raw bytes. This creates an API inconsistency and leaves the burden of proper decoding to the user.

According to the ZIP specification, a member's comment must be decoded using UTF-8 if its EFS flag bit is set. Otherwise, it should fallback to cp437 (per spec) or the local encoding (in practice), which typically aligns with the encoding used for filenames.

Currently, to properly read comments, users must write redundant boilerplate to manually check the EFS flag and apply the appropriate codec. This severely defeats the convenience introduced by the metadata_encoding parameter.

Proposed Solution
Prerequisite

For ZipInfo objects to handle comment encoding/decoding properly, there is a need to access the specified encoding context from the parent archive.

  • Approach A: Pass down the context as an internal attribute, such as _metadata_encoding. This aligns with the strategy used in PR gh-152846, which preserves the EFS flag during archive modification.
  • Approach B: Bind the parent ZipFile object via an internal attribute like _zipfile. This dynamically links the active encoding and enables future structural guardrails, such as detecting and warning against unsafe cross-archive metadata mutations, though care must be taken regarding pickling behavior.
Implementation

While updating {ZipFile,ZipInfo}.comment to be string-based would be the cleanest API, it is likely undesired due to breaking backward compatibility.

Alternatively, we can introduce a comment_text property (with getter, setter, and deleter) for both ZipFile and ZipInfo to manage string-based comments safely:

1. ZipFile.comment_text
  • get: Unlike member comments, the archive-level comment has no EFS flag in the ZIP specification. The getter should try decoding with UTF-8 first, then fallback to metadata_encoding (or 'cp437' if not provided).
  • set: Encodes with metadata_encoding (or 'cp437' if not provided). Falls back to UTF-8 if encoding fails, optionally issuing a warning.
  • delete: Clears the comment by resetting the underlying archive comment bytes to b''.
2. ZipInfo.comment_text
  • get: Takes the Unicode comment extra field (0x6375) if viable. Otherwise, decodes with UTF-8 if the EFS flag is set. Otherwise, uses the bound metadata_encoding (or 'cp437' if not provided).
  • set: Encodes with UTF-8 if the EFS flag is set. Otherwise, uses the bound metadata_encoding (or 'cp437' if not provided). If it fails, it can force the EFS flag and use UTF-8. An even simpler alternative is to always enforce EFS and UTF-8 when modifying via this property. Optionally clear or update the Unicode comment extra field.
  • delete: Clears the comment by resetting .comment to b''. Optionally clear the associated Unicode comment extra field.
Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia esaminando la gestione esistente dei commenti di ZipFile e ZipInfo e il parametro metadata_encoding. Determina se l'API debba associare il contesto di codifica o utilizzare un attributo di codifica interno, quindi definisci il comportamento compatibile di comment_text e i relativi casi limite. Il lavoro è completo quando il design scelto è implementato e sono presenti test per decodifica, codifica, eliminazione, gestione di EFS e comportamento di fallback.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
tooling
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Tranquilla
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.