creativecommons / creativecommons/quantifying

Add WikiCommons Data Source

Offen
#180 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
✨ goal: improvement 🏁 status: ready for work 💻 aspect: code help wanted 🟩 priority: low
Vorherrschende Sprache
Python
Sterne
48
Forks
74
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

## Problem
Hello, right now, the project collects data from Google Custom Search and GitHub, and work on adding Wikipedia is already in progress via PRs #176 and #167. Also, @TimidRobot commented about “more meaningful data” (for Wikipedia) suggesting they expect more than just basic counts — but WikiCommons they hasn’t been addressed yet.

However, WikiCommons is also an important source for Creative Commons–licensed media, and it’s not yet part of the automated system.
There’s an older version of it under `pre-automation/wikicommons/`, but it hasn’t been updated to the new structure.

## Description
Work can be done on adding WikiCommons as a new data source using the MediaWiki API.
This would collect counts of CC-licensed media files (like images, videos, and audio) by license type.

The plan is to:
- Review the old `pre-automation/wikicommons_scratcher.py` script.
- Rewrite it to match the new 3-step workflow (1-fetch, 2-process, 3-report).
- Make sure the new script uses the current shared helpers and output format.

This will help the project measure CC-licensed media content more accurately.

## Alternatives
It could be combined with the Wikipedia data, but keeping it separate makes it easier to track media content specifically.

## Additional context
- **Old script**: `pre-automation/wikicommons/`
- **API**:
- [MediaWiki Action API](https://commons.wikimedia.org/w/api.php)
- [API:Categorymembers - MediaWiki](https://www.mediawiki.org/wiki/API:Categorymembers/en)
- Builds on similar work done for **Wikipedia** (#159, #167, #176)

## Implementation
- [x] I would be interested in implementing this feature.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Lies zuerst pre-automation/wikicommons_scratcher.py und die zugehörigen Dateien in pre-automation/wikicommons/, und vergleiche dann die Wikipedia-Arbeiten in den Issues #159, #167 und #176. Passe die Quelle an den 1-fetch-, 2-process-, 3-report-Workflow unter Verwendung der aktuellen gemeinsamen Helfer und des aktuellen Ausgabeformats an; abgeschlossen ist die Aufgabe, wenn CC-lizenzierte Medien über die MediaWiki API nach Lizenztyp erfasst werden.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
data
Issue-Typ
Feature
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
50/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.