ArchiveBox / ArchiveBox/abx-plugins

New Extractor Idea: Internet Archive extractor

Open
#33 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
14
Forks
3
Avg merge
3h 12m
Merged PRs (30d)
15

Description

Extractor for Internet Archive content.

It should preserve the available item data and associated files across the different kinds of content hosted on archive.org, not just the rendered page.

This would make ArchiveBox treat Internet Archive URLs as archive collections with metadata and downloadable content, instead of only normal web pages.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the existing extractor and plugin implementations in this repository, then review how Internet Archive URLs and archive.org item data are represented. Define the supported content types and verify that the finished extractor preserves each item's available metadata and associated downloadable files rather than only a rendered page.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.