lucasew-graveyard / lucasew-graveyard/pdfnormalizer

Simplify bounding box logic

Open
#2 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Description

I found out that numpy has a simpler way do detect bounding box.

Ex:

![Captura de tela_2023-02-17_12-48-32](https://user-images.githubusercontent.com/15693688/219701367-bc1bc3b9-ddff-4af8-bc9d-6f6640a39b3d.png)

Contributor guide

No contributing guide indexed for this repository

Research direction

Locate the current bounding-box logic in the Python code and compare it with the NumPy approach shown in the issue. Confirm the simpler approach works for the project’s PDF extraction flow and update the relevant behavior without changing unrelated processing.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
computer-vision
Issue type
Refactor
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.