internetarchive / internetarchive/dweb-mirror

Crawl fail on some pages

Open
#325 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
JavaScript
Stars
330
Forks
34
PR merge metrics
No merged PRs in 30d

Description

Two problems here
a) crawl is failing to get [page](https://www-dweb-cors.dev.archive.org/BookReader/BookReaderImages.php?zip=%2F28%2Fitems%2Fkaputusan-kanda-pat-kawisesan%2Fkaputusan-kanda-pat-kawisesan-350ppi_jp2.zip&file=kaputusan-kanda-pat-kawisesan-350ppi_jp2%2Fkaputusan-kanda-pat-kawisesan-350ppi_0034.jp2&scale=4&rotate=0)
b) its saying no errors

Contributor guide

No contributing guide indexed for this repository

Research direction

Begin by reproducing the crawl failure for the linked BookReader page and record whether any error is emitted. Trace the crawl entry point and error-reporting path to identify why the page is not retrieved and why the failure is silent; done means the page crawls successfully or the failure is reported clearly.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.