camelot-dev / camelot-dev/excalibur

Unable to extract all rows from a pdf

Open
#171 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
238
PR merge metrics
No merged PRs in 30d

Description

Using Excalibur, my pdf contained 10 rows and only 5 rows have been extracted rest are simply blank.

Contributor guide

Open the contributing guide

Research direction

The issue does not identify a file, entry point, test, or sample PDF. Start by obtaining the PDF and reproducing the extraction in Excalibur, then trace the table-extraction path to determine why five of the ten rows are blank. Done means all ten rows are extracted correctly and the behavior is covered by a regression test.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.