ExtractPDF4J / ExtractPDF4J/ExtractPDF4J
Add ROI-based OCR extraction
- Dominant language
- Java
- Stars
- 524
- Forks
- 34
- PR merge metrics
- No merged PRs in 30d
Description
Perform OCR only on detected table regions.
---
## ✅ Challenge Spec
## 🧩 Problem
Explain the OCR / vision failure (noise, faint lines, rotation, low-res scans, etc.).
## 🎯 Goal
Improve extraction reliability/accuracy for scanned PDFs while keeping performance reasonable.
## ✅ Acceptance Criteria
- Demonstrable improvement on **at least 2** scanned sample PDFs/images (synthetic acceptable if shareable)
- Provide a debug overlay / screenshots / logs that show the improvement
- Add tests where feasible (or a reproducible sample corpus + run steps)
- Document any new config knobs (thresholding, OCR presets, ROI selection)
- No significant regression in extraction reliability
## 🧪 Validation
- Before/after evidence (images or extracted tables)
- Steps to reproduce locally
- Any performance impact noted
---
---
**Challenge Note:** This issue is part of the ExtractPDF4J Global Build Challenge 2026.
To participate: fork the repo → create a branch → submit a PR referencing this issue → Star the repo.
Contributor guide
Assessment
This issue has not been assessed yet.