coze-dev / coze-dev/coze-studio
Duplicate characters appear in extracted PDF headers/footers
Open
mod/knowledge
- Dominant language
- TypeScript
- Stars
- 21.6k
- Forks
- 3.1k
- PR merge metrics
- No merged PRs in 30d
Description
**Describe the bug**
After successfully uploading the PDF, the extracted text chunks contain duplicate characters.
Example:
Org:"ABC",
Extracted: "AABBCC"
, which are in headers and footers of PDFs.
The situation only appears when extracting files from specific website: [the_website](https://www.mohurd.gov.cn/?medium=01)
**Screenshots**
Contributor guide
Assessment
This issue has not been assessed yet.