CatchTheTornado / CatchTheTornado/text-extract-api
[feat] support JSON output chunking
- Dominant language
- Python
- Stars
- 3.2k
- Forks
- 279
- PR merge metrics
- No merged PRs in 30d
Description
The Marker OCR supports now pretty nice JSON output chunking including document structure: https://github.com/VikParuchuri/marker we could support it. The same thing is for `docling` (#54 ) -> https://ds4sd.github.io/docling/usage/#chunking
Contributor guide
No contributing guide indexed for this repository
Research direction
Review Marker’s linked JSON chunking support and Docling’s chunking guidance, then locate how this API currently exposes structured JSON output. The issue names no files or tests, so first clarify the supported behavior and define completion around validated chunked output for representative documents.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- api, backend
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100