codelucas / codelucas/newspaper
Sub-header/summary extraction
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 15.2k
- Forks
- 2.1k
- PR merge metrics
- No merged PRs in 30d
Description
Many articles have a sub-header or a summary; below the title but above the main text body.
As an example, the sub-header for this article https://www.fintechbusiness.com/industry/1185-bian-launches-api-market would be "Technology giant BIAN has officially launched its Open API platform at the SIBOS conference held in Sydney.".
Is there any way to (generically; not site-specific) extract that with newspaper?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Use the fintechbusiness.com example URL to reproduce the article extraction and compare the result with the quoted sub-header. Trace newspaper's existing article-text and metadata extraction flow, then define a generic way to expose that sub-header and verify it works without site-specific rules.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100