roboflow / roboflow/inference

`DocTR` model output missing important information that model produces

Open
#419 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
2.5k
Forks
319
Avg merge
1d 14h
Merged PRs (30d)
133

Description

Search before asking
  • I have searched the Inference issues and found no similar feature requests.
Description

DocTR produces not only text, but also location of the text. Our implementation linearises all of text data, skipping its location and not providing blocks-of-texts info.

We should think about:

  • making results of OCR models more generic (and apply that for inference server)
  • making changes into workflows blocks with OCR models to reflect those
Use case

No response

Additional

No response

Are you willing to submit a PR?
  • Yes I'd like to help by submitting a PR!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by tracing how DocTR OCR results are represented in the inference server and how workflows blocks consume OCR model output. Compare the text currently exposed with the location and block information DocTR produces. Done means the intended generic OCR result shape and corresponding workflows changes are defined and implemented, with coverage for the additional information.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend-api-design, computer-vision
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.