Script to archive all `related_urls` in archive.org
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 35/100
Research direction
Start by locating the dataset records containing related_urls and review the savepagenow interface for archiving pages in the Wayback Machine. The work is done when a script gathers every related_urls value and saves the URLs to archive.org.
Written by the indexing model from the issue text.
Description
Create a script that gathers all urls in related_urls fields and saves in archive.org in the Wayback Machine. Maybe using savepagenow.
- Dominant language
- Python
- Stars
- 486
- Forks
- 290
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from pyvideo/data
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Add PyCon US 2026 Open
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
PyCon SE 2019 Openevent::conference minimal_download
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
event::conference minimal_download
Difficulty 1/5 Under an hour Newbie friendliness 62/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100