camelot-dev / camelot-dev/excalibur

Running extraction across different pages

Open
#138 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
238
PR merge metrics
No merged PRs in 30d

Description

Is there a way to duplicate the rules (columns / table areas) across pages to check that the rules work? If the pages are slightly out of alignment, currently the extraction process doesn't work. For instance, I was looking to extract the information from pages 4,5,6 from the attached file. I add the columns for the first page and can't seem to copy these onto the second and 3rd pages.
[FTSE ASEAN Singapore copy 3_rep1.pdf](https://github.com/camelot-dev/excalibur/files/6322226/FTSE.ASEAN.Singapore.copy.3_rep1.pdf)

Contributor guide

Open the contributing guide

Research direction

Reproduce the request with the attached FTSE ASEAN Singapore PDF and pages 4–6. No source files, tests, or entry points are named in the issue, so first locate the extraction workflow and its handling of columns or table areas. Done should mean rules can be reused across slightly misaligned pages and extraction succeeds for the selected pages.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.