camelot-dev / camelot-dev/excalibur
Unable to extract all rows from a pdf
Open
- Dominant language
- Python
- Stars
- 1.8k
- Forks
- 238
- PR merge metrics
- No merged PRs in 30d
Description
Using Excalibur, my pdf contained 10 rows and only 5 rows have been extracted rest are simply blank.
Contributor guide
Research direction
The issue does not identify a file, entry point, test, or sample PDF. Start by obtaining the PDF and reproducing the extraction in Excalibur, then trace the table-extraction path to determine why five of the ten rows are blank. Done means all ten rows are extracted correctly and the behavior is covered by a regression test.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100