python / python/cpython

LZMA filtering documentation very poor

Offen
#121,186 0 Kommentare 1 Reaktion 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

docs
Vorherrschende Sprache
Python
Sterne
77.2k
Forks
35.9k
PR-Merge-Kennzahlen
PR-Kennzahlen ausstehend

Beschreibung

I'm trying to optimise compressing some documents, so I'm using LZMA's filters. These are really badly documented: https://docs.python.org/3/library/lzma.html#filter-chain-specs

General problems:

  • The filters are using short-hand for variable names lc, lp, mf, which is really bad practice for a publically available interface (and even for internal ones!). (ok, so this isn't a doc problem, but a library one).
  • It's not at all clear what any of the variables do. There's no links to any documentation (even third-party) which explain what any of these things do.
  • It's not clear what values are acceptable for some of them.
    • preset is documented elsewhere, and I assume the value range is the same?
    • lc doesn't say the maximum is 4, you have to read lp to see the combined maximum is 4, and then deduce.
    • lc, lp, pc, nice_len, don't state minimums
  • No defaults stated other than depth (0),
  • It's not at all clear what the different Filters do.
  • The limited blurb that does exist about Filter chains seems to assume a high degree of expertise in these things. For example: The last filter in the chain must be a compression filter, and any other filters must be delta or BCJ filters. - that has cleared nothing up.
  • What's a MODE_FAST or MODE_NORMAL do?
  • What's the difference between MF_HC3, MF_HC4, MF_BT2, MF_BT3, or MF_BT4?
  • What does lzma.PRESET_EXTREME do? There's no documentation for lzma.PRESET_DEFAULT at all (which my IDE exposed to me via autocomplete).

Etc etc.

Some specific problems as well:

  • 'filters': [{'id': 33, 'preset': 6, 'dict_size': 262144, 'lc': 1, 'lp': 1, 'mode': 2, 'nice_len': 1, 'mf': 3}]
    That creates an LZMAError. It's only after generating thousands of combinations of these that I notice a pattern: nice_len: 1. It seems there's a minimum nice_len size. This is not documented.

  • 'filters': [{'id': 33, 'preset': 2147483653, 'dict_size': 1073741824, 'lc': 1, 'lp': 0, 'mode': 2, 'nice_len': 201, 'mf': 20}]
    Generates a MemoryError (as do lots of others of my tests). No idea why as I have plenty of RAM. I presume some combination of those settings are causing an issue, but there's no documentation warning about this.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne mit dem Abschnitt zu den Spezifikationen von LZMA-Filterketten auf der verlinkten Python-Dokumentationsseite und überprüfe die Beschreibungen der Filteroptionen, Presets, Modi und Filterketten. Dokumentiere die Bedeutung der Parameter, gültige Wertebereiche, Standardwerte, die Unterschiede zwischen den Filtern sowie relevante Fehler- oder Speicherbeschränkungen, damit die Beispiele und Fehlerfälle erklärt werden.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
documentation
Issue-Typ
Dokumentation
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.