LAION-AI / LAION-AI/Open-Assistant

Proposal: Dataset based on subtitles for Japanese movies/tv shows from opensubtitles.org

Open
#2,747 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

data
Dominant language
Python
Stars
37.4k
Forks
3.3k
PR merge metrics
No merged PRs in 30d

Description

I just have finished a dataset very similar to https://github.com/LAION-AI/Open-Assistant/tree/main/data/datasets/fd_dialogue but for Japanese and taking the data from opensubtitles.org. The dataset contains subtitles for over 7000 tv shows and movies. The dataset is not formatted in the same style as the TV dialogue, with the speaker mentioned first, as that data is practically non-existing for Japanese, at least for the data in OpenSubitles. I was wondering if I should go on and do the pull request following the steps of the readme. Is this dataset good enough? The dataset is already uploaded on huggingface https://huggingface.co/datasets/Nan-Do/OpenSubtitlesJapanese

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the repository README contribution steps and inspect the linked Hugging Face OpenSubtitlesJapanese dataset, including its format and relationship to the existing fd_dialogue dataset. Done means maintainers have confirmed that the dataset is suitable and specified whether and how it should be contributed.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
18/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.