creativecommons / creativecommons/quantifying
Implement Internet Archive Data Fetching Pipeline
- Langage dominant
- Python
- Étoiles
- 48
- Forks
- 74
- Métriques de merge des PR
- Aucune PR mergée en 30 j
Description
## Problem
We currently lack a mechanism for systematically and regularly harvesting license-aware metadata from the Internet Archive (IA), a massive source of CC-licensed content. This data is essential for understanding the large-scale adoption and usage of Creative Commons licenses on the platform.
## Proposal / Solution
Implement a new Python script and pipeline specifically designed to query the Internet Archive's API, process the results, and generate aggregated statistics.
## Scope of Work (What this pipeline will do)
1. Query: Search the IA for all items with a Creative Commons/open-source license.
2. Data Fetching: Implement robust API calling with exponential backoff and retry logic to handle network and rate-limiting issues across large result sets.
3. Normalization: Clean and standardize the messy raw license URLs provided by the IA using a lookup table (ia_license_mapping.csv).
4. Aggregation: Collect counts for licenses, languages, and countries.
5. Output: Save the aggregated counts into three separate, version-controlled CSV files.
6. Automation: Include optional flags (--enable-save, --enable-git) to support dry runs and automated Git commits/pushes for regular data updates.
## Expected Outcome
We have a baseline and a recurring process for measuring CC usage on the Internet Archive.
## Related Links
- [Tools and APIs — Internet Archive Developer Portal](https://archive.org/developers/index-apis.html)
- [Internet Archive: Search Engine](https://archive.org/advancedsearch.php)
- [jjjake/internetarchive](https://github.com/jjjake/internetarchive): _A Python and Command-Line Interface to Archive.org_
- [Searching Items — Developer Interface — Internet Archive Developer Portal](https://archive.org/developers/internetarchive/api.html#searching-items)
- Pre-automation effort:
- [`pre-automation/internetarchive/`](https://github.com/creativecommons/quantifying/tree/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation/internetarchive)
- [`pre-automation/sources.md`](https://github.com/creativecommons/quantifying/blob/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation/sources.md)
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par lire pre-automation/internetarchive/ et pre-automation/sources.md, puis consultez la documentation liée de Internet Archive API et inspectez ia_license_mapping.csv. Le travail est considéré comme terminé lorsqu’un pipeline Python récurrent interroge et normalise les métadonnées de licence de IA, agrège les décomptes de licences, de langues et de pays dans trois CSVs versionnés, et prend en charge les flags --enable-save et --enable-git.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- data-engineering
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- À l'abandon
- Clarté
- Plutôt claire
- Accessibilité débutants
- 35/100