CatchTheTornado

CatchTheTornado/text-extract-api

View on GitHub

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

Stars
3.2k
Forks
279
Open beginner issues
0
Indexed issues
46
Dominant language
Python
License
MIT
Last GitHub push
Dec 8, 2025
Latest indexed
Sep 14, 2026
Contributing guide
No contributing guide
Code of conduct
No code of conduct
Beginner labels
help wanted good first issue
PR merge metrics
No merged PRs in 30d
46 open issues indexed Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.