camelot-dev / camelot-dev/excalibur

Data structure wrong matched (Example inside)

Open
#145 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
238
PR merge metrics
No merged PRs in 30d

Description

[Beispiel.pdf](https://github.com/camelot-dev/excalibur/files/7117387/Beispiel.pdf)

How can I trim the data recognition to get correct results from the example PDF provided?

1. Every first line has a number in the second (right) column
2. Text entries in the first (left) column might span up to three lines
3. Each page has to table areas (2-column-page-layout)

I appreciate evey help. It is some academic work for which I need the raw data...

Kind regards!

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the extraction behavior with the linked Beispiel.pdf and compare the output with the three layout rules in the issue. Done means the raw data correctly handles the right-column first-line numbers, left-column entries spanning up to three lines, and both table areas on each page.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.