aboutcode-org / aboutcode-org/typecode

PDF file detected as non-binary

未关闭
#41 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
10
派生
15
PR 合并指标
30 天内没有已合并 PR

描述

I am feeding a PDF file to `typecode.contenttype.is_binary`. As PDF files are usually considered as binary files, I would have expected the file to be detected as binary, but apparently the first bytes used for detection are looking like plain-text, leading to a wrong classification.

Example file: [antartica-3427135_640_1_libtiff.pdf](https://github.com/user-attachments/files/16411602/antartica-3427135_640_1_libtiff.pdf)

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。