creativecommons / creativecommons/quantifying

Implement Internet Archive Data Fetching Pipeline

Ouverte
#196 8 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
✨ goal: improvement 🏁 status: ready for work 💻 aspect: code help wanted 🟩 priority: low
Langage dominant
Python
Étoiles
48
Forks
74
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

## Problem

We currently lack a mechanism for systematically and regularly harvesting license-aware metadata from the Internet Archive (IA), a massive source of CC-licensed content. This data is essential for understanding the large-scale adoption and usage of Creative Commons licenses on the platform.

## Proposal / Solution
Implement a new Python script and pipeline specifically designed to query the Internet Archive's API, process the results, and generate aggregated statistics.

## Scope of Work (What this pipeline will do)
1. Query: Search the IA for all items with a Creative Commons/open-source license.
2. Data Fetching: Implement robust API calling with exponential backoff and retry logic to handle network and rate-limiting issues across large result sets.
3. Normalization: Clean and standardize the messy raw license URLs provided by the IA using a lookup table (ia_license_mapping.csv).
4. Aggregation: Collect counts for licenses, languages, and countries.
5. Output: Save the aggregated counts into three separate, version-controlled CSV files.
6. Automation: Include optional flags (--enable-save, --enable-git) to support dry runs and automated Git commits/pushes for regular data updates.

## Expected Outcome
We have a baseline and a recurring process for measuring CC usage on the Internet Archive.

## Related Links
- [Tools and APIs — Internet Archive Developer Portal](https://archive.org/developers/index-apis.html)
- [Internet Archive: Search Engine](https://archive.org/advancedsearch.php)
- [jjjake/internetarchive](https://github.com/jjjake/internetarchive): _A Python and Command-Line Interface to Archive.org_
- [Searching Items — Developer Interface — Internet Archive Developer Portal](https://archive.org/developers/internetarchive/api.html#searching-items)
- Pre-automation effort:
- [`pre-automation/internetarchive/`](https://github.com/creativecommons/quantifying/tree/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation/internetarchive)
- [`pre-automation/sources.md`](https://github.com/creativecommons/quantifying/blob/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation/sources.md)

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Commencez par lire pre-automation/internetarchive/ et pre-automation/sources.md, puis consultez la documentation liée de Internet Archive API et inspectez ia_license_mapping.csv. Le travail est considéré comme terminé lorsqu’un pipeline Python récurrent interroge et normalise les métadonnées de licence de IA, agrège les décomptes de licences, de langues et de pays dans trois CSVs versionnés, et prend en charge les flags --enable-save et --enable-git.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
data-engineering
Type d'issue
Fonctionnalité
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.