NanmiCoder / NanmiCoder/MediaCrawler

怎么选定时间段进行微博数据的爬取

Open
#501 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
65.3k
Forks
12.6k
PR merge metrics
No merged PRs in 30d

Description

在config文件中没有找到相关时间参数的配置,这个是需要修改代码逻辑嘛,如果需要应该在哪个文件中实现呀,非常感谢您开源代码的贡献!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review the config file and the Weibo crawler entry point to determine where a crawl time range could be configured and applied. Check the existing Weibo scraping flow first; the work is done when the supported time range is documented, configurable, and verified against the crawler's output.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.