OCR portuguese (PT-BR) document
- Dominant language
- Python
- Stars
- 9k
- Forks
- 775
- PR merge metrics
- No merged PRs in 30d
Description
[rgf_anexo_6_-_simplificado_-_2quadri_2023_-_executivo_1695992622.pdf](https://github.com/user-attachments/files/21017251/rgf_anexo_6_-_simplificado_-_2quadri_2023_-_executivo_1695992622.pdf)
In a simple PDF file, the return markdown brings Chinese characters and HTML code.
The document is a fiscal management report for a Brazilian municipality and is basically composed of tables with descriptions and values.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the OCR result with the attached Brazilian Portuguese PDF and inspect how the returned markdown represents its tables, characters, and HTML. Trace the document-parsing entry point that handles this PDF and compare the output with the Portuguese descriptions and values visible in the source document. Done means the document no longer produces Chinese characters or unintended HTML in the returned markdown.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- computer-vision
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100