feyninc / feyninc/pulpie

The same for .pdf

Open
#20 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
HTML
Stars
78
Forks
3
PR merge metrics
No merged PRs in 30d

Description

Often the problem with .pdf that parts of it (important text) are not grouped or structured. It would be nice to see this model being applied to such files to reconstruct a logically connected, grouped and structured text. I think that this would be a big deal.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by locating the existing model and extraction flow, then define the PDF inputs, reconstruction behavior, and tests needed to show that important text is logically connected, grouped, and structured.

Written by the indexing model from the issue text.

Assessment

Domain
content
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.