bigscience-workshop / bigscience-workshop/data_tooling

Create dataset vicon_visim400

Đang mở
#126 4 bình luận 0 reaction 1 người được giao Được @albertvillanova nhận Xem trên GitHub
data catalog need data sourcing feedback
Ngôn ngữ chính
HTML
Star
91
Fork
47
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

- uid: vicon_visim400
- type: processed
- description:
- name: Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness (ViCon and ViSim-400)
- description: This dataset consists of two kinds of datasets: The first dataset, namely ViCon, comprises pairs of synonyms and antonymys across noun, verb, and adjective classes, offerring data to distinguish between similarity and dissimilarity. The second dataset ViSim-400 is a dataset of semantic relation pairs which contains degrees of similarity across five semantic relations, as rated by human judges.
- homepage: https://www.ims.uni-stuttgart.de/forschung/ressourcen/experiment-daten/vnese-sem-datasets/
- validated: True
- languages:
- language_names:
- Vietnamese
- language_comments:
- language_locations:
- Western Europe
- Germany
- validated: False
- custodian:
- name: Thang Vu
- in_catalogue:
- type: A university or research institution
- location: Germany
- contact_name: Thang Vu
- contact_email: ngoc-thang.vu@ims.uni-stuttgart.de
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - it has a direct download link or links
- download_url: https://www.ims.uni-stuttgart.de/documents/ressourcen/experiment-daten/ViData.zip
- download_email:
- licensing:
- has_licenses: Unclear
- license_text:
- license_properties:
- license_list:
- pii:
- has_pii: No
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- processed_from_primary:
- from_primary: Original data
- primary_availability:
- primary_license:
- primary_types:
- validated: False
- media:
- category:
- text
- text_format:
- audiovisual_format:
- image_format:
- database_format:
- .ZIP
- text_is_transcribed: No
- instance_type: article
- instance_count: 100

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.