mimetypes raises UnicodeDecodeError when map files are not unicode encoded
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 77.2k
- 派生
- 36k
- PR 合并指标
- PR 指标待抓取
描述
Bug report
Bug description:
Hi, when I use the mimetypes module and one of the known mime.types files include a non utf-8 encoded comment, the operation fails with UnicodeDecodeError:
......
File "/usr/lib/python3.9/urllib/request.py", line 1506, in open_local_file
mtype = mimetypes.guess_type(filename)[0]
File "/usr/lib/python3.9/mimetypes.py", line 289, in guess_type
init()
File "/usr/lib/python3.9/mimetypes.py", line 362, in init
db.read(file)
File "/usr/lib/python3.9/mimetypes.py", line 204, in read
self.readfp(fp, strict)
File "/usr/lib/python3.9/mimetypes.py", line 215, in readfp
line = fp.readline()
File "/usr/lib/python3.9/codecs.py", line 322, in decode
(result, consumed) = self._buffer_decode(data, self.errors, final)
UnicodeDecodeError: 'utf-8' codec can't decode byte 0x83 in position 168: invalid start byte
The same can be forced with:
import mimetypes
mimetypes.init(files=["mimefile"])
and occurs because the file is opened in text mode expecting unicode encoding: https://github.com/python/cpython/blob/2e098abf95fc50f0d9a747140f8e17c5d8b65bce/Lib/mimetypes.py#L215
I am not sure whether there is a convention for which encoding the mime.types file will use, but I feel that at least comments should be allowed in any encoding?
CPython versions tested on:
3.9, 3.11
Operating systems tested on:
Linux, Other
Linked PRs
- gh-151216
- gh-155183
- gh-155184
- gh-155185
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 Lib/mimetypes.py 中 traceback 引用的 readfp() 路径,以及 mimetypes.init(files=["mimefile"]) 使用的文件打开行为开始。使用包含非 UTF-8 注释的 mime.types 文件重现该故障,然后验证此类注释不再导致 UnicodeDecodeError,同时 MIME 映射仍然可读。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- tooling
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 35/100