python / python/cpython

LZMA filtering documentation very poor

未关闭
#121,186 0 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

docs
主要语言
Python
星标
77.2k
派生
35.9k
PR 合并指标
PR 指标待抓取

描述

I'm trying to optimise compressing some documents, so I'm using LZMA's filters. These are really badly documented: https://docs.python.org/3/library/lzma.html#filter-chain-specs

General problems:

  • The filters are using short-hand for variable names lc, lp, mf, which is really bad practice for a publically available interface (and even for internal ones!). (ok, so this isn't a doc problem, but a library one).
  • It's not at all clear what any of the variables do. There's no links to any documentation (even third-party) which explain what any of these things do.
  • It's not clear what values are acceptable for some of them.
    • preset is documented elsewhere, and I assume the value range is the same?
    • lc doesn't say the maximum is 4, you have to read lp to see the combined maximum is 4, and then deduce.
    • lc, lp, pc, nice_len, don't state minimums
  • No defaults stated other than depth (0),
  • It's not at all clear what the different Filters do.
  • The limited blurb that does exist about Filter chains seems to assume a high degree of expertise in these things. For example: The last filter in the chain must be a compression filter, and any other filters must be delta or BCJ filters. - that has cleared nothing up.
  • What's a MODE_FAST or MODE_NORMAL do?
  • What's the difference between MF_HC3, MF_HC4, MF_BT2, MF_BT3, or MF_BT4?
  • What does lzma.PRESET_EXTREME do? There's no documentation for lzma.PRESET_DEFAULT at all (which my IDE exposed to me via autocomplete).

Etc etc.

Some specific problems as well:

  • 'filters': [{'id': 33, 'preset': 6, 'dict_size': 262144, 'lc': 1, 'lp': 1, 'mode': 2, 'nice_len': 1, 'mf': 3}]
    That creates an LZMAError. It's only after generating thousands of combinations of these that I notice a pattern: nice_len: 1. It seems there's a minimum nice_len size. This is not documented.

  • 'filters': [{'id': 33, 'preset': 2147483653, 'dict_size': 1073741824, 'lc': 1, 'lp': 0, 'mode': 2, 'nice_len': 201, 'mf': 20}]
    Generates a MemoryError (as do lots of others of my tests). No idea why as I have plenty of RAM. I presume some combination of those settings are causing an issue, but there's no documentation warning about this.

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

从链接的 Python 文档页面中的 LZMA 过滤器链规范部分开始,检查过滤器选项、预设、模式和过滤器链的说明。记录参数含义、有效范围、默认值、过滤器之间的差异以及相关的错误或内存限制,以便解释示例和失败情况。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
documentation
Issue 类型
文档
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。