python / python/cpython

tarfile: don't interpret GNU-style atime as ustar-style path prefix

Aberta
#155,629 1 comentário 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

stdlib type-bug
Linguagem predominante
Python
Estrelas
77.2k
Forks
36k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

Bug report

Bug description:

The GNU and ustar tar format families have slightly overlapping header definitions: ustar says that bytes 345:500 are used for the "path prefix" i.e. for paths longer than the 100 char limit in the path field, while old-style GNU ("oldgnu") says that 345:356 are used for storing an atime attribute for the member.

Consequently, the two overlap (partially), and a tar parser should check the header's magic to determine how to interpret that byte range.

At the moment, CPython does not use the header's magic, and instead unconditionally interprets that range as a ustar-style prefix:

https://github.com/python/cpython/blob/b11e749f7590e9a0907db908fa3e7e76c772c28f/Lib/tarfile.py#L1354

and then unconditionally uses that prefix as long as the member type (not the header type) isn't a special GNU member type:

https://github.com/python/cpython/blob/b11e749f7590e9a0907db908fa3e7e76c772c28f/Lib/tarfile.py#L1383-L1385

The end result of this is that tarfile can extract a file with a surprising name, whereas other parsers extract with the correct (non-ustar-prefixed) name.

MRE:

import io
import tarfile

member = tarfile.TarInfo("victim")
header = bytearray(member.tobuf(format=tarfile.GNU_FORMAT))

# Old-GNU atime field: valid octal timestamp 1.
header[345:357] = b"00000000001\0"

# Recalculate checksum.
header[148:156] = b" " * 8
header[148:156] = f"{sum(header):06o}\0 ".encode("ascii")

archive = bytes(header) + b"\0" * 1024

with tarfile.open(fileobj=io.BytesIO(archive), mode="r:") as tf:
    print(tf.getnames())

On a main build as of b11e749f7590e9a0907db908fa3e7e76c772c28f, this produces:

['00000000001/victim']

whereas the output should be ['victim'], since the format is GNU_FORMAT instead of a ustar-family format.

I think the fix for this is to tweak the obj.name assignment to only use prefix when the magic bytes match POSIX_MAGIC, i.e. not GNU_MAGIC or any legacy (v7, pre-ustar) magic.

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Linked PRs
  • gh-155706

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece em Lib/tarfile.py, na análise do prefixo e na atribuição de obj.name próximas às linhas referenciadas, e depois reproduza o problema com o MRE fornecido. Está concluído quando o arquivo GNU-format retornar ['victim'] em vez de ['00000000001/victim']; verifique o PR vinculado gh-155706 antes de começar.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
tooling
Tipo de issue
Bug
Dificuldade
2/5
Tempo estimado
1-3 horas
Status de atividade
Estagnada
Clareza
Claramente especificada
Facilidade para iniciantes
35/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.