coze-dev / coze-dev/coze-studio

Duplicate characters appear in extracted PDF headers/footers

Open
#1,788 0 comments 0 reactions 1 assignee Claimed by @liuyunchao-1998 View on GitHub
mod/knowledge
Dominant language
TypeScript
Stars
21.6k
Forks
3.1k
PR merge metrics
No merged PRs in 30d

Description

**Describe the bug**

After successfully uploading the PDF, the extracted text chunks contain duplicate characters.
Example:
Org:"ABC",
Extracted: "AABBCC"
, which are in headers and footers of PDFs.

The situation only appears when extracting files from specific website: [the_website](https://www.mohurd.gov.cn/?medium=01)

**Screenshots**

Image

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.