bytedance / bytedance/Dolphin

OCR portuguese (PT-BR) document

Open
#96 1 comment 4 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
9k
Forks
775
PR merge metrics
No merged PRs in 30d

Description

[rgf_anexo_6_-_simplificado_-_2quadri_2023_-_executivo_1695992622.pdf](https://github.com/user-attachments/files/21017251/rgf_anexo_6_-_simplificado_-_2quadri_2023_-_executivo_1695992622.pdf)
In a simple PDF file, the return markdown brings Chinese characters and HTML code.
The document is a fiscal management report for a Brazilian municipality and is basically composed of tables with descriptions and values.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the OCR result with the attached Brazilian Portuguese PDF and inspect how the returned markdown represents its tables, characters, and HTML. Trace the document-parsing entry point that handles this PDF and compare the output with the Portuguese descriptions and values visible in the source document. Done means the document no longer produces Chinese characters or unintended HTML in the returned markdown.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.