creativecommons / creativecommons/quantifying
[Meta] Ways to Contribute
- Dominant language
- Python
- Stars
- 48
- Forks
- 74
- PR merge metrics
- No merged PRs in 30d
Description
## Overview
⚠️ **This meta-issue should not be worked on itself.** Instead it is a place for me (@TimidRobot) to document ways to engage with this project.
## First
- Read the documentation on [Welcome — Creative Commons Open Source](https://opensource.creativecommons.org/)
- **[Contribution Guidelines — Creative Commons Open Source](https://opensource.creativecommons.org/contributing-code/)** 👀
## Ways to Contribute
- Pull Requests (PRs)
- Contribute a PR for an existing issue per **[Contribution Guidelines — Creative Commons Open Source](https://opensource.creativecommons.org/contributing-code/)**
- Help review PRs
- Issues
- Create a new issue recommending a data source that should be included. Include information like:
- quantity of records
- types of metadata available
- API documentation link
- API requirements and limitations
- Also see
- [`sources.md`](https://github.com/creativecommons/quantifying/blob/main/sources.md)
- [Sources | Openverse](https://openverse.org/sources) (any of the listed sources are potential sources for this project)
- [`pre-automation/`](https://github.com/creativecommons/quantifying/tree/925e7212e98ffe0ee166439c10ae951cf41304dc/pre-automation)
- Create a new issue related to a single script or data source:
- Scripts should be using `.env` ([theskumar/python-dotenv](https://github.com/theskumar/python-dotenv)) and not `query_secrets.py` or similar
- Scripts mustn't be monolithic--they should be limited to a single phase (ex. query, process, report. See #22)
- Scripts must be designed to be run from the repository root via pipenv (ex. `pipenv run PATH/SCRIPT.PY`)
- Script should determine its own path and set appropriate global variables (ex. `DIR_ROOT`, `DIR_SCRIPT`)
- Scripts have a lot of duplication between them. Begin a shared library (remember to keep issues as small and descrete as possible--limit each issue/PR to a single script or data source).
- Scripts should be using retries with exponential backoff (ex. #2)
## Tips
### Conventions and best practices
- Always sort data and lists (both implicit and explicit) naturally ([Natural sort order - Wikipedia](https://en.wikipedia.org/wiki/Natural_sort_order))
- Example: sorted constants
https://github.com/creativecommons/quantifying/blob/46dcd3ff9d66e172b87d79b097adcb9b348c6f2c/scripts/1-fetch/wikipedia_fetch.py#L32-L44
- Example: sorted data
```python
from operator import itemgetter
```
```python
data.sort(key=itemgetter("TOOL_IDENTIFIER", "CATEGORY_CODE"))
```
- #2
- #217
### Plot process data
Any plots in phase 3-report should graph data without significantly modifying it. This means that development of phase 2-process and phase 3-report usually needs to be done at the same time.
Put another way, there should be a 1:1 relationship between phase 3-report plots and phase 2-process CSV files.
Contributor guide
Assessment
This issue has not been assessed yet.