pdf2json Performance over large PDF
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 2.2k
- Forks
- 394
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 7
Description
Hi All,
I have a PDF file that contains about 500 pages (3.6mb) - I can't post because it contains sensitive data. When I load it up through pdf2json, it takes about 10 minutes to fire the dataReady callback... is this expected?
I am running the node application on an macbook pro, i7, 16GB... and seriously expected it to be faster.
The PDF contents are of a timetable nature... and all I want to extract are the text strings and their x/y locations for grouped by page.
Does anyone else have performance issues with pdf2json... or does anyone else have any suggestions as to other node modules to use for this purpose?
Looking forward to some help... and free to answer any questions.
Ta.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the report with a comparable 500-page PDF and the requested text and x/y extraction, measuring the time before the dataReady callback. Read the pdf2json parsing path while comparing results on smaller files; done means identifying whether the delay is expected and documenting or addressing a confirmed performance problem.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- node.js
- Domain
- backend, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100