creativecommons / creativecommons/quantifying

[Feature] Post-GSoC '24: Solidify Processing Scripts for Quarterly Analysis

Đang mở
#124 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
✨ goal: improvement 🏁 status: ready for work 💬 talk: discussion 💻 aspect: code 🟩 priority: low
Ngôn ngữ chính
Python
Star
48
Fork
74
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

## Context

Automating Quantifying the Commons was a project endeavor for the Google Summer of Code 2024 program, in which a baseline automation software for data gathering, processing, and analysis was successfully developed. However given the time and resource constraints that we had to consider, there are still addressable endeavors to improve this codebase over the upcoming quarters and years. This is the first (1) of five (5) issues raised specifically for post-GSoC contributions.

## Problem

Due to only having one quarter’s worth of data, current processing scripts (`2-process`) are not fully optimized for long-term data analysis, which makes it difficult to accurately assess trends and patterns over quarterly periods.

## Description

This feature involves refining the processing scripts to handle data collected over a larger period, enabling more robust quarterly analysis. The focus will be on adding code that can effectively compare details of each data source by quarter (ex. `2024Q3` data is compared to all previous quarters’ data) and adding them into separate datasets for report generation.

**NOTE**: since contributing to this specific issue is limited by access to API data fetching and the fact that the solution is long-term, this issue is being set as a discussion for all open-source developers to be able to pitch their ideas for final implementation by the developer(s) who work on the codebase.

## Implementation

- [x] I would be interested in implementing this feature.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu bằng cách đọc các script xử lý trong `2-process` và theo dõi cách dữ liệu được lấy qua API được thu thập và biểu diễn qua các quý. Kết quả dự kiến là các tập dữ liệu riêng biệt so sánh từng nguồn dữ liệu, chẳng hạn như 2024Q3, với tất cả các quý trước đó để tạo báo cáo; các chi tiết triển khai vẫn để ngỏ để thảo luận.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
data-engineering
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.