bigscience-workshop / bigscience-workshop/data_tooling

Create dataset australian_twittersphere

未关闭 适合新手
#133 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
data catalog need custodian permission
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

描述

- uid: australian_twittersphere
- type: processed
- description:
- name: Australian Twittersphere
- description: The Australian Twittersphere is a longitudinal, curated collection of tweets from approximately 838,000 Twitter accounts identified as ‘Australian’. The Digital Observatory has maintained reliable, ongoing data collection since early 2018, with approximately 23 million tweets being collected per month. There is also an archive of approximately 2 billion tweets from 2006 to 2016. The Digital Observatory currently collects approximately 37 million tweets per month.
- homepage: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: True
- languages:
- language_names:
- English
- language_comments: Australian English
- language_locations:
- Oceania
- Australia
- validated: False
- custodian:
- name: Digital Observatory of the Queensland University of Technology
- in_catalogue:
- type: A university or research institution
- location: Australia
- contact_name:
- contact_email: digitalobservatory@qut.edu.au
- contact_submitter: False
- additional: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory
- validated: False
- availability:
- procurement:
- for_download: No - we would need to spontaneously reach out to the current owners/custodians
- download_url:
- download_email: https://www.qut.edu.au/research/why-qut/infrastructure/digital-observatory/services-and-equipment
- licensing:
- has_licenses: Unclear
- license_text: The data should be able to be used to train models while respecting the rights and wishes of the data creators and custodians, as they were obtained in compliance with Twitter's terms of use.
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: other
- no_pii_justification_text: The data were obtained from Twitter and should have been anonimysed.
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: Yes - their documentation/homepage/description is available
- primary_license: Yes - the dataset curators have obtained consent from the source material owners
- primary_types:
- web | social media
- validated: False
- from_primary_entries:
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type: post
- instance_count: n>1B
- instance_size: 10

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先检查数据集条目的格式,并将提供的 Australian Twittersphere 记录添加到 australian_twittersphere.json 中。检查字段、值和文件名是否符合周围目录的约定;当新的数据集条目通过 repository 的验证流程后,工作即完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
json
领域
data
Issue 类型
功能
难度
1/5
预计耗时
1 小时以内
活跃度
停滞
描述清晰度
描述清楚
新手友好度
65/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。