python / python/cpython

tarfile: unbounded memory use on large pax and GNU extensions

Ouverte
#155,633 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

stdlib type-bug
Langage dominant
Python
Étoiles
77.2k
Forks
35.9k
Métriques de merge des PR
Métriques de PR en attente

Description

Bug report

Bug description:

This is somewhat similar to https://github.com/python/cpython/issues/151497, and fits under the larger umbrella of https://github.com/python/cpython/issues/141713.

Summary

Both the pax and GNU tar families support extensions, via different mechanisms. These extensions have pre-declared lengths, and a tar parser must read a payload of length bytes to consume them.

Prior to https://github.com/python/cpython/issues/151497 this was done in a single read(n) call, resulting in a single large up-front allocation. That was changed to _safe_read with https://github.com/python/cpython/pull/151498, which bounds each read call to 1MB.

This prevents unbounded memory consumption at the read site, but not in aggregate. For example, an attacker can still contrive a pax-style tar archive with an extremely large individual pax record, and tarfile will buffer that pax record (in 1MB increments) into memory. The same is true for GNU extensions.

Solution

I think the solution is to put a reasonable caps on the sizes of extensions.

This could be done at a few different layers (e.g. restricting individual pax record sizes versus the entire pax extension size), but I think doing it at the extension size layer is probably simplest and most consistent.

My proposal would be:

  1. No pax or GNU extension should ever exceed 1 MB in raw size (i.e., the size reported by its tar frame). This is extremely conservative, i.e. should be well above what any real-world tar would need to put in its extensions.
  2. For pax in particular, the global pax extension state should never exceed some reasonable multiplier of the extension cap. For example, someone shouldn't be able to induce higher memory usage by chaining g -> g -> g -> ... -> file.txt, where each g member has 1MB of pax extension state.

For prior art, we perform this kind of bounding in tar-codec, e.g. here:

https://github.com/astral-sh/tar-codec/blob/dbd4b5efeb6edb732c993d107a7cfdc83f8d29a8/crates/tar-framing/src/stream.rs#L1099-L1154

and we impose a default cap of 256KB for pax extensions, 1MB for all active global pax extensions, and 128KB for GNU extensions:

https://github.com/astral-sh/tar-codec/blob/dbd4b5efeb6edb732c993d107a7cfdc83f8d29a8/crates/tar-framing/src/lib.rs#L124-L137

(These numbers are not particularly scientific; we picked them because we think even 256KB is very conservative i.e. high for pax, and 128KB for GNU is well beyond what any normal OS will accept as a pathname length limit.)

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par l’analyse des extensions pax et GNU de tarfile, y compris le chemin _safe_read existant, et examinez les tests associés. Reproduisez des archives contenant des extensions surdimensionnées ou chaînées, puis définissez et testez des limites pour la taille brute des extensions et l’état pax global accumulé afin que l’analyse ne puisse pas faire croître la mémoire sans limite.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
security
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Calme
Clarté
Plutôt claire
Accessibilité débutants
48/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.