bigscience-workshop / bigscience-workshop/data_tooling

Create dataset australian_twittersphere

Ouverte Adaptée aux débutants
#133 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
data catalog need custodian permission
Langage dominant
HTML
Étoiles
91
Forks
47
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

- uid: australian_twittersphere
- type: processed
- description:
- name: Australian Twittersphere
- description: The Australian Twittersphere is a longitudinal, curated collection of tweets from approximately 838,000 Twitter accounts identified as ‘Australian’. The Digital Observatory has maintained reliable, ongoing data collection since early 2018, with approximately 23 million tweets being collected per month. There is also an archive of approximately 2 billion tweets from 2006 to 2016. The Digital Observatory currently collects approximately 37 million tweets per month.
- homepage: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: True
- languages:
- language_names:
- English
- language_comments: Australian English
- language_locations:
- Oceania
- Australia
- validated: False
- custodian:
- name: Digital Observatory of the Queensland University of Technology
- in_catalogue:
- type: A university or research institution
- location: Australia
- contact_name:
- contact_email: digitalobservatory@qut.edu.au
- contact_submitter: False
- additional: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: False
- availability:
- procurement:
- for_download: No - we would need to spontaneously reach out to the current owners/custodians
- download_url:
- download_email: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory/services-and-equipment
- licensing:
- has_licenses: Unclear
- license_text: The data should be able to be used to train models while respecting the rights and wishes of the data creators and custodians, as they were obtained in compliance with Twitter's terms of use.
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: other
- no_pii_justification_text: The data were obtained from Twitter and should have been anonimysed.
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - their documentation/homepage/description is available
- primary_license: Yes - the dataset curators have obtained consent from the source material owners
- primary_types:
- web | social media
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type: post
- instance_count: n>1B
- instance_size: 10

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Start by inspecting the dataset entry format and add the supplied Australian Twittersphere record to australian_twittersphere.json. Check that the fields, values, and filename match the surrounding catalog conventions; the work is done when the new dataset entry is accepted by the repository's validation process.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
json
Domaine
data
Type d'issue
Fonctionnalité
Difficulté
1/5
Temps estimé
Moins d'une heure
Activité
À l'abandon
Clarté
Clairement spécifiée
Accessibilité débutants
65/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.