codeforboston / codeforboston/maple

Remove Line Numbers from Bill PDF Scraping

Open
#2,157 3 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
TypeScript
Stars
56
Forks
175
Avg merge
2d 5h
Merged PRs (30d)
13

Description

## Summary

We just added a new codepath to scrape the Document text of a "Bill" from the legislature-provided PDF as a fallback for when the legislature doesn't make that text available via the API's `content.DocumentText` field.

It looks like this is working reasonably well, but we've just noticed an issue in one particular case: in addition to the text of the bill, we are also scraping the line numbers (which is technically legible, but looks noticeably off).

To remedy this, we should filter out line number when scraping the text from a bill's PDF.

## Success Criteria
- [ ] PDF-scraped text should also filter out line numbers
- [ ] Existing successful text scraping should be unaffected

## Additional Links
- Example bill that hit this issue: https://maple-dev.vercel.app/bills/194/H5469
- This has a `DocumentText` of `null` in the API, so it will use the PDF fallback, but we have also scraped line numbers (visible if you go to the page and click "View Text" to open the bill text modal).

@Smoss Related to your recent work on PDF scraping

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.