creativecommons / creativecommons/quantifying

[Feature] Post-GSoC '24: Solidify Processing Scripts for Quarterly Analysis

Offen
#124 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
✨ goal: improvement 🏁 status: ready for work 💬 talk: discussion 💻 aspect: code 🟩 priority: low
Vorherrschende Sprache
Python
Sterne
48
Forks
74
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

## Context

Automating Quantifying the Commons was a project endeavor for the Google Summer of Code 2024 program, in which a baseline automation software for data gathering, processing, and analysis was successfully developed. However given the time and resource constraints that we had to consider, there are still addressable endeavors to improve this codebase over the upcoming quarters and years. This is the first (1) of five (5) issues raised specifically for post-GSoC contributions.

## Problem

Due to only having one quarter’s worth of data, current processing scripts (`2-process`) are not fully optimized for long-term data analysis, which makes it difficult to accurately assess trends and patterns over quarterly periods.

## Description

This feature involves refining the processing scripts to handle data collected over a larger period, enabling more robust quarterly analysis. The focus will be on adding code that can effectively compare details of each data source by quarter (ex. `2024Q3` data is compared to all previous quarters’ data) and adding them into separate datasets for report generation.

**NOTE**: since contributing to this specific issue is limited by access to API data fetching and the fact that the solution is long-term, this issue is being set as a discussion for all open-source developers to be able to pitch their ideas for final implementation by the developer(s) who work on the codebase.

## Implementation

- [x] I would be interested in implementing this feature.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by reading the processing scripts under `2-process` and tracing how API-fetched data is gathered and represented across quarters. The intended result is separate datasets that compare each data source, such as 2024Q3, with all previous quarters for report generation; implementation details remain open for discussion.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
data-engineering
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.