The same for .pdf
Open
- Dominant language
- HTML
- Stars
- 78
- Forks
- 3
- PR merge metrics
- No merged PRs in 30d
Description
Often the problem with .pdf that parts of it (important text) are not grouped or structured. It would be nice to see this model being applied to such files to reconstruct a logically connected, grouped and structured text. I think that this would be a big deal.
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are named. Start by locating the existing model and extraction flow, then define the PDF inputs, reconstruction behavior, and tests needed to show that important text is logically connected, grouped, and structured.
Written by the indexing model from the issue text.
Assessment
- Domain
- content
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100