ExtractPDF4J / ExtractPDF4J/ExtractPDF4J

Add ROI-based OCR extraction

Open
#76 0 comments 0 reactions 0 assignees View on GitHub
advanced challenge-march-2026 challenge-track-ocr
Dominant language
Java
Stars
524
Forks
34
PR merge metrics
No merged PRs in 30d

Description

Perform OCR only on detected table regions.

---

## ✅ Challenge Spec

## 🧩 Problem
Explain the OCR / vision failure (noise, faint lines, rotation, low-res scans, etc.).

## 🎯 Goal
Improve extraction reliability/accuracy for scanned PDFs while keeping performance reasonable.

## ✅ Acceptance Criteria
- Demonstrable improvement on **at least 2** scanned sample PDFs/images (synthetic acceptable if shareable)
- Provide a debug overlay / screenshots / logs that show the improvement
- Add tests where feasible (or a reproducible sample corpus + run steps)
- Document any new config knobs (thresholding, OCR presets, ROI selection)
- No significant regression in extraction reliability

## 🧪 Validation
- Before/after evidence (images or extracted tables)
- Steps to reproduce locally
- Any performance impact noted

---

---

**Challenge Note:** This issue is part of the ExtractPDF4J Global Build Challenge 2026.
To participate: fork the repo → create a branch → submit a PR referencing this issue → Star the repo.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.