getRawTextContent() returns text not found in pdfData.formImage
Nobody has claimed this yet.
- Dominant language
- Java
- Stars
- 2.2k
- Forks
- 394
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 7
Description
When I look at the pdfData.formImage object there are no text blocks, but getRawTextContent() returns the strings that I expect to see. I'm not very familiar with how/if the implementation of text extraction differs, so is there a simple explanation for this e.g. the formatting of the pdf, or how do I go about debugging?
pdfData.formImage is something like
{"formImage":{"Transcoder":"pdf2json@1.2.0 [https://github.com/modesty/pdf2json]","Agency":"","Id":{"AgencyId":"","Name":"","MC":false,"Max":1,"Parent":""},"Pages":[{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":52.625,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.5,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.036,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":49.036,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":47.932,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":47.932,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]},{"Height":47.767,"HLines":[],"VLines":[],"Fills":[{"x":0,"y":0,"w":0,"h":0,"clr":1}],"Texts":[],"Fields":[],"Boxsets":[]}],"Width":38.25}}
```
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing getRawTextContent() and comparing its output with pdfData.formImage.Pages[].Texts in the supplied example. Investigate the PDF text-extraction path and document why text can be returned when those arrays are empty; done when the behavior has a clear explanation or a reproducible correction is identified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- json
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100