CatchTheTornado / CatchTheTornado/text-extract-api

Investigate and test pdf-extract-kit models

Open
#32 0 comments 0 reactions 0 assignees View on GitHub
help wanted
Dominant language
Python
Stars
3.2k
Forks
279
PR merge metrics
No merged PRs in 30d

Description

https://github.com/opendatalab/PDF-Extract-Kit?tab=readme-ov-file

Contributor guide

No contributing guide indexed for this repository

Research direction

No repository files, tests, or entry points are named. Start with the linked PDF-Extract-Kit README, then determine how its models could be investigated and tested in text-extract-api. Done is not defined by the issue; the scope, target models, integration point, and test criteria need to be agreed before implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.