bigscience-workshop / bigscience-workshop/data_tooling
Create dataset UIT-ViHSD
- Ngôn ngữ chính
- HTML
- Star
- 91
- Fork
- 47
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Mô tả
- uid: UIT-ViHSD
- type: processed
- description:
- name: Vietnamese Hate Speech Detection Dataset
- description: In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social media, hate speech has become a critical problem for social network users. To solve this problem, we introduce the ViHSD – a human-annotated dataset for automatically detecting hate speech on the social network. This dataset contains over 30,000 comments, each comment in the dataset has one of three labels: CLEAN, OFFENSIVE, or HATE. Besides, we introduce the data creation process for annotating and evaluating the quality of the dataset. Finally, we evaluated the dataset by deep learning models and transformer models.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects#h.fs21gpd5w6p1
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments:
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: Mr. Son Luu
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Son Luu
- contact_email: sonlt@uit.edu.vn
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: No - but the current owners/custodians have contact information for data queries
- download_url:
- download_email: sonlt@uit.edu.vn
- licensing:
- has_licenses: Unclear
- license_text:
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count:
- instance_size:
- validated: False
- fname: UIT-ViHSD.json
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
Bắt đầu với mục nhập UIT-ViHSD.json được yêu cầu và xác minh rằng siêu dữ liệu của mục nhập khớp với các chi tiết dataset được cung cấp, bao gồm ngôn ngữ tiếng Việt, danh mục văn bản, thông tin liên hệ của đơn vị quản lý, tính khả dụng và các trường xác thực. Được xem là hoàn tất khi bản ghi dataset được thêm vào với tên tệp được chỉ định và tất cả các giá trị được cung cấp được giữ nguyên.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Lĩnh vực
- data
- Loại issue
- Tính năng
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức độ hoạt động
- Đình trệ
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 50/100