creativecommons / creativecommons/quantifying
[Feature] Post-GSoC '24: Solidify Processing Scripts for Quarterly Analysis
- Vorherrschende Sprache
- Python
- Sterne
- 48
- Forks
- 74
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
## Context
Automating Quantifying the Commons was a project endeavor for the Google Summer of Code 2024 program, in which a baseline automation software for data gathering, processing, and analysis was successfully developed. However given the time and resource constraints that we had to consider, there are still addressable endeavors to improve this codebase over the upcoming quarters and years. This is the first (1) of five (5) issues raised specifically for post-GSoC contributions.
## Problem
Due to only having one quarter’s worth of data, current processing scripts (`2-process`) are not fully optimized for long-term data analysis, which makes it difficult to accurately assess trends and patterns over quarterly periods.
## Description
This feature involves refining the processing scripts to handle data collected over a larger period, enabling more robust quarterly analysis. The focus will be on adding code that can effectively compare details of each data source by quarter (ex. `2024Q3` data is compared to all previous quarters’ data) and adding them into separate datasets for report generation.
**NOTE**: since contributing to this specific issue is limited by access to API data fetching and the fact that the solution is long-term, this issue is being set as a discussion for all open-source developers to be able to pitch their ideas for final implementation by the developer(s) who work on the codebase.
## Implementation
- [x] I would be interested in implementing this feature.
Beitragsleitfaden
Rechercherichtung
Start by reading the processing scripts under `2-process` and tracing how API-fetched data is gathered and represented across quarters. The intended result is separate datasets that compare each data source, such as 2024Q3, with all previous quarters for report generation; implementation details remain open for discussion.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- data-engineering
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100