LZMA filtering documentation very poor
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 35.9k
- PR 合并指标
- PR 指标待抓取
描述
I'm trying to optimise compressing some documents, so I'm using LZMA's filters. These are really badly documented: https://docs.python.org/3/library/lzma.html#filter-chain-specs
General problems:
- The filters are using short-hand for variable names
lc,lp,mf, which is really bad practice for a publically available interface (and even for internal ones!). (ok, so this isn't a doc problem, but a library one). - It's not at all clear what any of the variables do. There's no links to any documentation (even third-party) which explain what any of these things do.
- It's not clear what values are acceptable for some of them.
-
presetis documented elsewhere, and I assume the value range is the same?
-
lcdoesn't say the maximum is 4, you have to readlpto see the combined maximum is 4, and then deduce.
-
lc,lp,pc,nice_len, don't state minimums
- No defaults stated other than
depth (0), - It's not at all clear what the different Filters do.
- The limited blurb that does exist about Filter chains seems to assume a high degree of expertise in these things. For example:
The last filter in the chain must be a compression filter, and any other filters must be delta or BCJ filters.- that has cleared nothing up. - What's a
MODE_FASTorMODE_NORMALdo? - What's the difference between
MF_HC3, MF_HC4, MF_BT2, MF_BT3, or MF_BT4? - What does
lzma.PRESET_EXTREMEdo? There's no documentation forlzma.PRESET_DEFAULTat all (which my IDE exposed to me via autocomplete).
Etc etc.
Some specific problems as well:
-
'filters': [{'id': 33, 'preset': 6, 'dict_size': 262144, 'lc': 1, 'lp': 1, 'mode': 2, 'nice_len': 1, 'mf': 3}]
That creates anLZMAError. It's only after generating thousands of combinations of these that I notice a pattern:nice_len: 1. It seems there's a minimum nice_len size. This is not documented. -
'filters': [{'id': 33, 'preset': 2147483653, 'dict_size': 1073741824, 'lc': 1, 'lp': 0, 'mode': 2, 'nice_len': 201, 'mf': 20}]
Generates aMemoryError(as do lots of others of my tests). No idea why as I have plenty of RAM. I presume some combination of those settings are causing an issue, but there's no documentation warning about this.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从链接的 Python 文档页面中的 LZMA 过滤器链规范部分开始,检查过滤器选项、预设、模式和过滤器链的说明。记录参数含义、有效范围、默认值、过滤器之间的差异以及相关的错误或内存限制,以便解释示例和失败情况。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- documentation
- Issue 类型
- 文档
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 35/100