scrapinghub / scrapinghub/dateparser
Date ambiguities during search_dates
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.9k
- Forks
- 520
- Avg merge
- 22h 56m
- Merged PRs (30d)
- 6
Description
For the following
from dateparser.search import search_dates
result_date = search_dates("Statement: 4 February 2019 10 4")
print(result_date)
We get the result:
[('4 February', datetime.datetime(2020, 2, 4, 0, 0)), ('2019 10', datetime.datetime(2019, 10, 4, 0, 0))]
The last numbers are not dates, but are seen as such. That is okay because "2019 10 4" could make a valid date. But then "4 February 2019" should still be identified correctly.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the example through the dateparser.search.search_dates entry point and inspect how the text is split into date candidates. The fix is done when “4 February 2019” is identified as one date without incorrectly treating the trailing numbers as part of separate dates, while the reported input still parses correctly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 42/100