python / python/cpython

LZMA filtering documentation very poor

Aberta
#121,186 0 comentários 1 reação 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

docs
Linguagem predominante
Python
Estrelas
77.2k
Forks
36k
Métricas de merge de PRs
Métricas de PR pendentes

Descrição

I'm trying to optimise compressing some documents, so I'm using LZMA's filters. These are really badly documented: https://docs.python.org/3/library/lzma.html#filter-chain-specs

General problems:

  • The filters are using short-hand for variable names lc, lp, mf, which is really bad practice for a publically available interface (and even for internal ones!). (ok, so this isn't a doc problem, but a library one).
  • It's not at all clear what any of the variables do. There's no links to any documentation (even third-party) which explain what any of these things do.
  • It's not clear what values are acceptable for some of them.
    • preset is documented elsewhere, and I assume the value range is the same?
    • lc doesn't say the maximum is 4, you have to read lp to see the combined maximum is 4, and then deduce.
    • lc, lp, pc, nice_len, don't state minimums
  • No defaults stated other than depth (0),
  • It's not at all clear what the different Filters do.
  • The limited blurb that does exist about Filter chains seems to assume a high degree of expertise in these things. For example: The last filter in the chain must be a compression filter, and any other filters must be delta or BCJ filters. - that has cleared nothing up.
  • What's a MODE_FAST or MODE_NORMAL do?
  • What's the difference between MF_HC3, MF_HC4, MF_BT2, MF_BT3, or MF_BT4?
  • What does lzma.PRESET_EXTREME do? There's no documentation for lzma.PRESET_DEFAULT at all (which my IDE exposed to me via autocomplete).

Etc etc.

Some specific problems as well:

  • 'filters': [{'id': 33, 'preset': 6, 'dict_size': 262144, 'lc': 1, 'lp': 1, 'mode': 2, 'nice_len': 1, 'mf': 3}]
    That creates an LZMAError. It's only after generating thousands of combinations of these that I notice a pattern: nice_len: 1. It seems there's a minimum nice_len size. This is not documented.

  • 'filters': [{'id': 33, 'preset': 2147483653, 'dict_size': 1073741824, 'lc': 1, 'lp': 0, 'mode': 2, 'nice_len': 201, 'mf': 20}]
    Generates a MemoryError (as do lots of others of my tests). No idea why as I have plenty of RAM. I presume some combination of those settings are causing an issue, but there's no documentation warning about this.

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Direção de pesquisa

Comece pela seção de especificações de cadeias de filtros LZMA na página vinculada da documentação do Python e revise as descrições das opções de filtro, dos presets, dos modos e das cadeias de filtros. Documente o significado dos parâmetros, os intervalos válidos, os valores padrão, as diferenças entre os filtros e as restrições relevantes de erro ou memória, para que os exemplos e os casos de falha sejam explicados.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
python
Domínio
documentation
Tipo de issue
Documentação
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Estagnada
Clareza
Razoavelmente clara
Facilidade para iniciantes
35/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.