camelot-dev / camelot-dev/excalibur

Unstructured Data

Open
#137 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
238
PR merge metrics
No merged PRs in 30d

Description

Hi team, the camelot and excalibur is a great tool for extracting data from pdf but sometimes I get unstructured data.
Please give me some suggestion or a way to handle this type of problem below is the attachment you can see 
![pdftable](https://user-images.githubusercontent.com/73991817/107609972-68769d00-6c66-11eb-8c33-bdefbc24df9c.JPG)
![xlsxfile](https://user-images.githubusercontent.com/73991817/107610011-91972d80-6c66-11eb-9e3b-aaf46af17fb9.JPG)
so here the instrument type is **nestle india** and industry type is **consumer non durables** it takes the **Durables** as an extra cell
Please i request you to provide me some solution to overcome this problem.

Thank you so much for making this library and tool.

Contributor guide

Open the contributing guide

Research direction

No source file, test, or entry point is named. Start by examining the attached screenshots and reproducing the extraction result with the same PDF; determine why “Durables” becomes an extra cell, and consider the issue done when the resulting table preserves the shown instrument and industry values.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.