python / python/cpython

tarfile: don't interpret GNU-style atime as ustar-style path prefix

Ouverte
#155,629 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

stdlib type-bug
Langage dominant
Python
Étoiles
77.2k
Forks
36k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

The GNU and ustar tar format families have slightly overlapping header definitions: ustar says that bytes 345:500 are used for the "path prefix" i.e. for paths longer than the 100 char limit in the path field, while old-style GNU ("oldgnu") says that 345:356 are used for storing an atime attribute for the member.

Consequently, the two overlap (partially), and a tar parser should check the header's magic to determine how to interpret that byte range.

At the moment, CPython does not use the header's magic, and instead unconditionally interprets that range as a ustar-style prefix:

https://github.com/python/cpython/blob/b11e749f7590e9a0907db908fa3e7e76c772c28f/Lib/tarfile.py#L1354

and then unconditionally uses that prefix as long as the member type (not the header type) isn't a special GNU member type:

https://github.com/python/cpython/blob/b11e749f7590e9a0907db908fa3e7e76c772c28f/Lib/tarfile.py#L1383-L1385

The end result of this is that tarfile can extract a file with a surprising name, whereas other parsers extract with the correct (non-ustar-prefixed) name.

MRE:

import io
import tarfile

member = tarfile.TarInfo("victim")
header = bytearray(member.tobuf(format=tarfile.GNU_FORMAT))

# Old-GNU atime field: valid octal timestamp 1.
header[345:357] = b"00000000001\0"

# Recalculate checksum.
header[148:156] = b" " * 8
header[148:156] = f"{sum(header):06o}\0 ".encode("ascii")

archive = bytes(header) + b"\0" * 1024

with tarfile.open(fileobj=io.BytesIO(archive), mode="r:") as tf:
    print(tf.getnames())

On a main build as of b11e749f7590e9a0907db908fa3e7e76c772c28f, this produces:

['00000000001/victim']

whereas the output should be ['victim'], since the format is GNU_FORMAT instead of a ustar-family format.

I think the fix for this is to tweak the obj.name assignment to only use prefix when the magic bytes match POSIX_MAGIC, i.e. not GNU_MAGIC or any legacy (v7, pre-ustar) magic.

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Linked PRs
  • gh-155706

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans Lib/tarfile.py, au niveau de l’analyse du préfixe et de l’affectation de obj.name autour des lignes référencées, puis reproduisez le problème avec le MRE fourni. C’est terminé lorsque l’archive GNU-format renvoie ['victim'] au lieu de ['00000000001/victim'] ; vérifiez le PR lié gh-155706 avant de commencer.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
tooling
Type d'issue
Bug
Difficulté
2/5
Temps estimé
1-3 heures
Activité
À l'abandon
Clarté
Clairement spécifiée
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.