creativecommons / creativecommons/quantifying

[Feature] Post-GSoC '24: Solidify Processing Scripts for Quarterly Analysis

Abierto
#124 1 comentario 0 reacciones 0 asignados Ver en GitHub
✨ goal: improvement 🏁 status: ready for work 💬 talk: discussion 💻 aspect: code 🟩 priority: low
Lenguaje dominante
Python
Estrellas
48
Forks
74
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

## Context

Automating Quantifying the Commons was a project endeavor for the Google Summer of Code 2024 program, in which a baseline automation software for data gathering, processing, and analysis was successfully developed. However given the time and resource constraints that we had to consider, there are still addressable endeavors to improve this codebase over the upcoming quarters and years. This is the first (1) of five (5) issues raised specifically for post-GSoC contributions.

## Problem

Due to only having one quarter’s worth of data, current processing scripts (`2-process`) are not fully optimized for long-term data analysis, which makes it difficult to accurately assess trends and patterns over quarterly periods.

## Description

This feature involves refining the processing scripts to handle data collected over a larger period, enabling more robust quarterly analysis. The focus will be on adding code that can effectively compare details of each data source by quarter (ex. `2024Q3` data is compared to all previous quarters’ data) and adding them into separate datasets for report generation.

**NOTE**: since contributing to this specific issue is limited by access to API data fetching and the fact that the solution is long-term, this issue is being set as a discussion for all open-source developers to be able to pitch their ideas for final implementation by the developer(s) who work on the codebase.

## Implementation

- [x] I would be interested in implementing this feature.

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Comienza leyendo los scripts de procesamiento bajo `2-process` y siguiendo cómo se recopilan y representan los datos obtenidos mediante la API a lo largo de los trimestres. El resultado previsto son conjuntos de datos separados que comparen cada fuente de datos, como 2024Q3, con todos los trimestres anteriores para la generación de informes; los detalles de implementación quedan abiertos a discusión.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
data-engineering
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.