`DocTR` model output missing important information that model produces
Open
Nobody has claimed this yet.
enhancement
- Dominant language
- Python
- Stars
- 2.5k
- Forks
- 319
- Avg merge
- 1d 14h
- Merged PRs (30d)
- 133
Description
Search before asking
- I have searched the Inference issues and found no similar feature requests.
Description
DocTR produces not only text, but also location of the text. Our implementation linearises all of text data, skipping its location and not providing blocks-of-texts info.
We should think about:
- making results of OCR models more generic (and apply that for
inferenceserver) - making changes into
workflowsblocks with OCR models to reflect those
Use case
No response
Additional
No response
Are you willing to submit a PR?
- Yes I'd like to help by submitting a PR!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing how DocTR OCR results are represented in the inference server and how workflows blocks consume OCR model output. Compare the text currently exposed with the location and block information DocTR produces. Done means the intended generic OCR result shape and corresponding workflows changes are defined and implemented, with coverage for the additional information.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend-api-design, computer-vision
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100