creativecommons / creativecommons/quantifying

[Meta] Ways to Contribute

Abierto
#39 14 comentarios 1 reacción 1 asignado Reclamado por @TimidRobot Ver en GitHub
✨ goal: improvement 🌟 goal: addition 🏷 status: label work required 💬 talk: discussion 💻 aspect: code 🤖 aspect: dx good first issue help wanted 🟨 priority: medium
Lenguaje dominante
Python
Estrellas
48
Forks
74
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

## Overview
⚠️ **This meta-issue should not be worked on itself.** Instead it is a place for me (@TimidRobot) to document ways to engage with this project.

## First
- Read the documentation on [Welcome — Creative Commons Open Source](https://opensource.creativecommons.org/)
- **[Contribution Guidelines — Creative Commons Open Source](https://opensource.creativecommons.org/contributing-code/)** 👀

## Ways to Contribute
- Pull Requests (PRs)
- Contribute a PR for an existing issue per **[Contribution Guidelines — Creative Commons Open Source](https://opensource.creativecommons.org/contributing-code/)**
- Help review PRs
- Issues
- Create a new issue recommending a data source that should be included. Include information like:
- quantity of records
- types of metadata available
- API documentation link
- API requirements and limitations
- Also see
- [`sources.md`](https://github.com/creativecommons/quantifying/blob/main/sources.md)
- [Sources | Openverse](https://openverse.org/sources) (any of the listed sources are potential sources for this project)
- [`pre-automation/`](https://github.com/creativecommons/quantifying/tree/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation)
- Create a new issue related to a single script or data source:
- Scripts should be using `.env` ([theskumar/python-dotenv](https://github.com/theskumar/python-dotenv)) and not `query_secrets.py` or similar
- Scripts mustn't be monolithic--they should be limited to a single phase (ex. query, process, report. See #22)
- Scripts must be designed to be run from the repository root via pipenv (ex. `pipenv run PATH/SCRIPT.PY`)
- Script should determine its own path and set appropriate global variables (ex. `DIR_ROOT`, `DIR_SCRIPT`)
- Scripts have a lot of duplication between them. Begin a shared library (remember to keep issues as small and descrete as possible--limit each issue/PR to a single script or data source).
- Scripts should be using retries with exponential backoff (ex. #2)

## Tips

### Conventions and best practices
- Always sort data and lists (both implicit and explicit) naturally ([Natural sort order - Wikipedia](https://en.wikipedia.org/wiki/Natural_sort_order))
- Example: sorted constants
https://github.com/creativecommons/quantifying/blob/46dcd3ff9d66e172b87d79b097adcb9b348c6f2c/scripts/1-fetch/wikipedia_fetch.py#L32-L44
- Example: sorted data
```python
from operator import itemgetter
```
```python
data.sort(key=itemgetter("TOOL_IDENTIFIER", "CATEGORY_CODE"))
```
- #2
- #217

### Plot process data
Any plots in phase 3-report should graph data without significantly modifying it. This means that development of phase 2-process and phase 3-report usually needs to be done at the same time.

Put another way, there should be a 1:1 relationship between phase 3-report plots and phase 2-process CSV files.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.