bigscience-workshop / bigscience-workshop/data_tooling

Create dataset tsac

Open
#352 0 comments 0 reactions 0 assignees View on GitHub
data catalog
Dominant language
HTML
Stars
91
Forks
47
PR merge metrics
No merged PRs in 30d

Description

- uid: tsac
- type: primary
- description:
- name: TSAC
- description: Tunisian sentiment analysis corpus. The dataset contains 17K user comments manually annotated to positive and negative. The corpus is collected from Facebook comments on official pages of Tunisian radios and TV channels. The dataset is collected between Jan 2015 and June 2016.
- homepage: https://github.com/fbougares/TSAC
- validated: True
- languages:
- language_names:
- Arabic
- ar-TN
- language_comments:
- language_locations:
- Northern Africa
- Tunisia
- validated: False
- custodian:
- name:
- in_catalogue:
- type:
- location:
- contact_name:
- contact_email:
- contact_submitter: False
- additional:
- validated: False
- availability:
- procurement:
- for_download: Yes - it has a direct download link or links
- download_url: https://github.com/fbougares/TSAC
- download_email:
- licensing:
- has_licenses: Yes
- license_text: Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.
- license_properties:
- open license
- license_list:
- lgpl-3.0: GNU Lesser General Public License v3.0 only
- pii:
- has_pii: Unclear
- generic_pii_likely:
- generic_pii_list:
- numeric_pii_likely:
- numeric_pii_list:
- sensitive_pii_likely:
- sensitive_pii_list:
- no_pii_justification_class: general knowledge not written by or referring to private persons
- no_pii_justification_text:
- validated: False
- source_category:
- category_type: website
- category_web:
- category_media:
- validated: False
- media:
- category:
- text
- text_format:
- .TXT
- audiovisual_format:
- image_format:
- database_format:
- .ZIP
- text_is_transcribed: No
- instance_type: post
- instance_count: 1K

Contributor guide

No contributing guide indexed for this repository

Research direction

Create the dataset entry in tsac.json using the supplied TSAC metadata, download URL, licensing details, language, and media fields. Review the issue text against the resulting file and confirm the entry is complete and consistently structured with the requested values.

Written by the indexing model from the issue text.

Assessment

Domain
data
Issue type
Feature
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.