bigscience-workshop / bigscience-workshop/data_tooling

Create dataset UIT-ViHSD

Đang mở
#123 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
data catalog need custodian permission
Ngôn ngữ chính
HTML
Star
91
Fork
47
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

- uid: UIT-ViHSD
- type: processed
- description:
- name: Vietnamese Hate Speech Detection Dataset
- description: In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social media, hate speech has become a critical problem for social network users. To solve this problem, we introduce the ViHSD – a human-annotated dataset for automatically detecting hate speech on the social network. This dataset contains over 30,000 comments, each comment in the dataset has one of three labels: CLEAN, OFFENSIVE, or HATE. Besides, we introduce the data creation process for annotating and evaluating the quality of the dataset. Finally, we evaluated the dataset by deep learning models and transformer models.
- homepage: https://sites.google.com/uit.edu.vn/uit-nlp/datasets-projects#h.fs21gpd5w6p1
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments:
- language_locations:
- South-eastern Asia
- Vietnam
- validated: False
- custodian:
- name: Mr. Son Luu
- in_catalogue:
- type: A university or research institution
- location: Vietnam
- contact_name: Mr. Son Luu
- contact_email: sonlt@uit.edu.vn
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: No - but the current owners/custodians have contact information for data queries
- download_url:
- download_email: sonlt@uit.edu.vn
- licensing:
- has_licenses: Unclear
- license_text:
- license_properties:
- license_list:
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Taken from primary source
- primary_availability: No - the dataset curators kept the source data secret
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- text_is_transcribed: No
- instance_type:
- instance_count:
- instance_size:
- validated: False
- fname: UIT-ViHSD.json

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu với mục nhập UIT-ViHSD.json được yêu cầu và xác minh rằng siêu dữ liệu của mục nhập khớp với các chi tiết dataset được cung cấp, bao gồm ngôn ngữ tiếng Việt, danh mục văn bản, thông tin liên hệ của đơn vị quản lý, tính khả dụng và các trường xác thực. Được xem là hoàn tất khi bản ghi dataset được thêm vào với tên tệp được chỉ định và tất cả các giá trị được cung cấp được giữ nguyên.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Lĩnh vực
data
Loại issue
Tính năng
Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
50/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.