internetarchive / internetarchive/dweb-mirror
Crawl fail on some pages
- Dominant language
- JavaScript
- Stars
- 330
- Forks
- 34
- PR merge metrics
- No merged PRs in 30d
Description
Two problems here
a) crawl is failing to get [page](https://www-dweb-cors.dev.archive.org/BookReader/BookReaderImages.php?zip=%2F28%2Fitems%2Fkaputusan-kanda-pat-kawisesan%2Fkaputusan-kanda-pat-kawisesan-350ppi_jp2.zip&file=kaputusan-kanda-pat-kawisesan-350ppi_jp2%2Fkaputusan-kanda-pat-kawisesan-350ppi_0034.jp2&scale=4&rotate=0)
b) its saying no errors
Contributor guide
No contributing guide indexed for this repository
Research direction
Begin by reproducing the crawl failure for the linked BookReader page and record whether any error is emitted. Trace the crawl entry point and error-reporting path to identify why the page is not retrieved and why the failure is silent; done means the page crawls successfully or the failure is reported clearly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- javascript
- Domain
- backend
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100