ExtractPDF4J / ExtractPDF4J/ExtractPDF4J

Improve handling of sparse tables

Open
#70 0 comments 0 reactions 0 assignees View on GitHub
challenge-march-2026 challenge-track-parser intermediate
Dominant language
Java
Stars
524
Forks
34
PR merge metrics
No merged PRs in 30d

Description

Improve robustness when tables contain empty cells.

---

## ✅ Challenge Spec

## 🧩 Problem
Describe the limitation / failure mode this issue addresses (include real-world context where possible).

## 🎯 Goal
State the expected improvement clearly (accuracy, stability, correctness, etc.).

## ✅ Acceptance Criteria
- Improvement demonstrated on **at least 2** representative sample PDFs (synthetic PDFs acceptable if shareable)
- Add/update **unit or integration tests** for the behaviour
- No regression in existing test suite / CI checks
- Update docs/Javadoc if behaviour or configuration changes

## 🧪 Validation
- Provide before/after output comparison (or screenshot/log excerpt)
- Link test evidence and sample inputs (or reproducible steps)

---

---

**Challenge Note:** This issue is part of the ExtractPDF4J Global Build Challenge 2026.
To participate: fork the repo → create a branch → submit a PR referencing this issue → Star the repo.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.