creativecommons / creativecommons/quantifying
Add Zenodo as Data Source for Commons Quantification
- Lenguaje dominante
- Python
- Estrellas
- 48
- Forks
- 74
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
## Problem
Zenodo is a major repository for open access research outputs with 5.5M+ records, but is not currently included in our commons quantification project. Adding Zenodo would significantly expand our coverage of Creative Commons licensed content, particularly in academic and research domains.
## Description
Implement data collection from Zenodo using their REST API to gather license information for quantifying the commons. This involves:
- Fetching records with structured license metadata
- Classifying Creative Commons and other open licenses
- Generating reports by year, resource type, and language
- Handling API rate limiting and pagination
## Zenodo Useful Links
### Official Documentation
- **REST API Documentation**: https://developers.zenodo.org/
- **API Reference**: https://zenodo.org/api/docs
- **General Documentation**: https://help.zenodo.org/
- **Developer Documentation**: https://developers.zenodo.org/
- **Search Guide**: https://help.zenodo.org/guides/search/
- **Zenodo Homepage**: https://zenodo.org/
### API Endpoints
- **Base URL**: `https://zenodo.org/api/records`
- **Records Search**: `https://zenodo.org/api/records`
- **Single Record**: `https://zenodo.org/api/records/{id}`
- **Communities**: `https://zenodo.org/api/communities`
## Technical Details
### Query Strategy
```
GET https://zenodo.org/api/records?q=*&size=100&page=1&sort=bestmatch
```
**Parameters:**
- `q`: Query string (use `*` for all records)
- `size`: Records per page (300) *implementation choice*
- `page`: Page number for pagination
- `sort`: Sorting method (bestmatch recommended)
### API Types Available
1. **REST API** (Recommended)
- Format: JSON
- Authentication: None required for public records
- Structured license data: `metadata.license.id`
2. **OAI-PMH** (Not recommended)
- Format: XML Dublin Core
- Unreliable license parsing from free-text fields `(dc:rights)`
### Key Metadata Fields
- **License**: `metadata.license.id` (structured, e.g., "cc-by-4.0")
- **Access Rights**: `metadata.access_right` ("open", "restricted", "embargoed")
- **Publication Date**: `metadata.publication_date` (ISO format)
- **Resource Type**: `metadata.resource_type.title`
- **Language**: `metadata.language` (ISO codes)
## Implementation
- [x] I would be interested in implementing this feature.
Guía de contribución
Línea de trabajo
Empieza con la documentación de la REST API de Zenodo y el endpoint /api/records, centrándote en la paginación, los límites de tasa y los campos de metadatos enumerados en el issue. Se considera terminado cuando se puedan recopilar y clasificar los registros por licencia, con informes que cubran el año, el tipo de recurso y el idioma, gestionando la paginación y los límites de tasa.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- data-engineering
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100