bigscience-workshop / bigscience-workshop/data_tooling

Create dataset UIT-ViHSD

未关闭
#123 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
data catalog need custodian permission
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

描述

- uid: UIT-ViHSD
- type: processed
- description:
- name: Vietnamese Hate Speech Detection Dataset
- description: In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social media, hate speech has become a critical problem for social network users. To solve this problem, we introduce the ViHSD – a human-annotated dataset for automatically detecting hate speech on the social network. This dataset contains over 30,000 comments, each comment in the dataset has one of three labels: CLEAN, OFFENSIVE, or HATE. Besides, we introduce the data creation process for annotating and evaluating the quality of the dataset. Finally, we evaluated the dataset by deep learning models and transformer models.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects#h.fs21gpd5w6p1
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments:
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: Mr. Son Luu
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Son Luu
- contact_email: sonlt@uit.edu.vn
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: No - but the current owners/custodians have contact information for data queries
- download_url:
- download_email: sonlt@uit.edu.vn
- licensing:
- has_licenses: Unclear
- license_text:
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count:
- instance_size:
- validated: False
- fname: UIT-ViHSD.json

贡献指南

这个仓库没有索引到贡献指南

调研方向

从请求的 UIT-ViHSD.json 条目开始,并验证其元数据是否与提供的数据集详细信息匹配,包括越南语、文本类别、维护者联系方式、可用性和验证字段。添加具有指定文件名的数据集记录,并保留所有提供的值,即表示完成。

由索引模型根据 Issue 内容生成。

评估

领域
data
Issue 类型
功能
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
基本清楚
新手友好度
50/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。