scrapinghub / scrapinghub/dateparser

Parsing date with mixed slashes and Chinese time of day

Open
#808 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Type: Bug - Language
Dominant language
Python
Stars
2.9k
Forks
520
Avg merge
22h 56m
Merged PRs (30d)
6

Description

similar to #232

datestring 2020/10/9 下午 02:26:26 is used in zh-CN language HFS servers
Other example: https://twitter.com/healthyscc/status/1299582251101949952

import dateparser

d = dateparser.parse("2020/10/9 下午 02:26:26")
print(d)

returns None

expect: datetime object

Thanks for the work on this library and making it open source

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the provided Python reproduction for 2020/10/9 下午 02:26:26 and inspect the Chinese-language date and time-of-day parsing paths, especially handling of mixed separators. Add a regression test for this input; done means parsing returns the expected datetime object without breaking existing Chinese date parsing.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
internationalization
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.