JakeYallop / JakeYallop/WaybackDownloader

Iterate snapshots in reverse

Open
#2 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

help wanted
Dominant language
C#
Stars
12
Forks
1
PR merge metrics
No merged PRs in 30d

Description

Generally, more-up-to-date versions of a page are on later snapshot pages. Each individual page is already iterated in reverse, but we could avoid much more downloading by iterating the complete list of pages in reverse as well. This would need to use showNumPages=true to get the maximum number of pages.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by locating the code that retrieves Wayback Machine snapshot pages and examine how it determines the page count. Use the showNumPages=true request to obtain the maximum page number, then iterate the complete page list in reverse so newer snapshot pages are downloaded first and unnecessary downloads are avoided.

Written by the indexing model from the issue text.

Assessment

Tech stack
csharp
Domain
cli
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.