creativecommons / creativecommons/quantifying

Add Zenodo as Data Source for Commons Quantification

Open
#249 0 comments 0 reactions 0 assignees View on GitHub
✨ goal: improvement 🏁 status: ready for work 💻 aspect: code help wanted 🟩 priority: low
Dominant language
Python
Stars
48
Forks
74
PR merge metrics
No merged PRs in 30d

Description

## Problem
Zenodo is a major repository for open access research outputs with 5.5M+ records, but is not currently included in our commons quantification project. Adding Zenodo would significantly expand our coverage of Creative Commons licensed content, particularly in academic and research domains.

## Description
Implement data collection from Zenodo using their REST API to gather license information for quantifying the commons. This involves:

- Fetching records with structured license metadata
- Classifying Creative Commons and other open licenses
- Generating reports by year, resource type, and language
- Handling API rate limiting and pagination

## Zenodo Useful Links

### Official Documentation
- **REST API Documentation**: https://developers.zenodo.org/
- **API Reference**: https://zenodo.org/api/docs
- **General Documentation**: https://help.zenodo.org/
- **Developer Documentation**: https://developers.zenodo.org/
- **Search Guide**: https://help.zenodo.org/guides/search/
- **Zenodo Homepage**: https://zenodo.org/

### API Endpoints
- **Base URL**: `https://zenodo.org/api/records`
- **Records Search**: `https://zenodo.org/api/records`
- **Single Record**: `https://zenodo.org/api/records/{id}`
- **Communities**: `https://zenodo.org/api/communities`

## Technical Details

### Query Strategy
```
GET https://zenodo.org/api/records?q=*&size=100&page=1&sort=bestmatch
```

**Parameters:**
- `q`: Query string (use `*` for all records)
- `size`: Records per page (300) *implementation choice*
- `page`: Page number for pagination
- `sort`: Sorting method (bestmatch recommended)

### API Types Available
1. **REST API** (Recommended)
- Format: JSON
- Authentication: None required for public records
- Structured license data: `metadata.license.id`

2. **OAI-PMH** (Not recommended)
- Format: XML Dublin Core
- Unreliable license parsing from free-text fields `(dc:rights)`

### Key Metadata Fields
- **License**: `metadata.license.id` (structured, e.g., "cc-by-4.0")
- **Access Rights**: `metadata.access_right` ("open", "restricted", "embargoed")
- **Publication Date**: `metadata.publication_date` (ISO format)
- **Resource Type**: `metadata.resource_type.title`
- **Language**: `metadata.language` (ISO codes)

## Implementation

- [x] I would be interested in implementing this feature.

Contributor guide

Open the contributing guide

Research direction

Start with Zenodo's REST API documentation and the /api/records endpoint, focusing on pagination, rate limits, and the metadata fields listed in the issue. Done means records can be collected and classified by license, with reports covering year, resource type, and language while handling pagination and rate limiting.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.