lm-sys / lm-sys/FastChat

Language distribution of ShareGPT 70K conversation dataset for FastChat T5

Open
#1,607 2 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

What are all the languages present in the ShareGPT 70,000 conversation dataset which was used to fine-tune FastChat-T5?

The ReadMe file points to data_cleaning.md which was used to get data from ShareGPT. Within data_cleaning.md seems like sharegpt_clean_lang.json contains the list of languages in consideration and some languages are skipped.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with docs/commands/data_cleaning.md and inspect sharegpt_clean_lang.json to understand which languages were considered or skipped. Compare that information with the ShareGPT 70,000-conversation dataset, then document the complete set of languages present and clarify the filtering used.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data, machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.