fani-lab / fani-lab/LADy

Dataset for LADy (Story Board)

Open
#72 22 comments 0 reactions 2 assignees Claimed by @hosseinfani View on GitHub
dataset
Dominant language
Python
Stars
7
Forks
7
PR merge metrics
No merged PRs in 30d

Description

**Literature Review:**

- Other datasets from resource tracks of cs conferences in text analysis domain

-- Sigir
-- Ecir
-- Cikm
-- …

- Other datasets in review analysis domain

-- SemEval
-- Amazon
-- …

**Our Dataset:**

- So far: Remove explicit aspect(s) (current approach) → grammar is destroyed

-- To show the grammar errors, we use llms or grammar check tools to show the % of errors
-- Show methods of aspect detection relying on the syntax tree

- Method 1: Remove explicit aspect(s) and rewrite the review → no grammar issue

-- double-check no errors with the same method above (llms or grammar check tools)

- Method 2: Originally latent

-- The human annotator should label the latent aspect(s), Double check through crowdsourcing to students, Voting on what’s the latent aspect(s)

**Review Domains**

-- Electronics,
-- Clothes,
-- …

**Dataset File Structure**

**Datasheet (data card) for datasets**

**Licensing**

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.