LAION-AI / LAION-AI/Open-Assistant
Proposal: Dataset based on subtitles for Japanese movies/tv shows from opensubtitles.org
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.4k
- Forks
- 3.3k
- PR merge metrics
- No merged PRs in 30d
Description
I just have finished a dataset very similar to https://github.com/LAION-AI/Open-Assistant/tree/main/data/datasets/fd_dialogue but for Japanese and taking the data from opensubtitles.org. The dataset contains subtitles for over 7000 tv shows and movies. The dataset is not formatted in the same style as the TV dialogue, with the speaker mentioned first, as that data is practically non-existing for Japanese, at least for the data in OpenSubitles. I was wondering if I should go on and do the pull request following the steps of the readme. Is this dataset good enough? The dataset is already uploaded on huggingface https://huggingface.co/datasets/Nan-Do/OpenSubtitlesJapanese
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the repository README contribution steps and inspect the linked Hugging Face OpenSubtitlesJapanese dataset, including its format and relationship to the existing fd_dialogue dataset. Done means maintainers have confirmed that the dataset is suitable and specified whether and how it should be contributed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 18/100