bigscience-workshop / bigscience-workshop/data_tooling

Create dataset global_voices_portuguese

Open
#153 0 comments 0 reactions 0 assignees View on GitHub
data catalog
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

- uid: global_voices_portuguese
- type: primary
- description:
- name: Global Voices Portuguese
- description: Global Voices pages in Portuguese
- homepage: https://pt.globalvoices.org/
- validated: True
- languages:
- language_names:
- Portuguese
- language_comments:
- language_locations:
- Americas
- Europe
- Western Africa
- Timor-Leste
- validated: False
- custodian:
- name:
- in_catalogue: global_voices
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - it has a direct download link or links
- download_url: https://pt.globalvoices.org/
- download_email:
- licensing:
- has_licenses: Yes
- license_text: https://globalvoices.org/about/global-voices-attribution-policy/
- license_properties:
- open license
- license_list:
- cc-by-3.0: Creative Commons Attribution 3.0 Unported
- pii:
- has_pii: Yes
- generic_pii_likely: very likely
- generic_pii_list:
- names
- email addresses
- numeric_pii_likely: somewhat likely
- numeric_pii_list:
- telephone numbers
- sensitive_pii_likely: very likely
- sensitive_pii_list:
- racial or ethnic origin
- political opinions
- religious or philosophical beliefs
- no_pii_justification_class:
- no_pii_justification_text:
- validated: False
- source_category:
- category_type: website
- category_web: news or magazine website
- category_media:
- validated: False
- media:
- category:
- text
- text_format:
- .HTML
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type: article
- instance_count: 1K

Contributor guide

No contributing guide indexed for this repository

Research direction

Create global_voices_portuguese.json from the complete dataset record in the issue. First locate the repository's existing dataset catalogue files and compare their structure, then add the new record and verify that its filename, fields, URLs, language, licensing, and privacy metadata match the issue.

Written by the indexing model from the issue text.

Assessment

Domain
data
Issue type
Feature
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.