facebookresearch / facebookresearch/SONAR
SONAR training dataset in Huggingface
- Dominant language
- Python
- Stars
- 910
- Forks
- 103
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
My goal is to train Sparse Autoencoders (SAEs). If the team can share the training dataset in huggingface, then it would be possible to train SAEs that learn features from the training dataset
The problem:
1. To train SAEs, you need to use the dataset that was used to train SONAR (i.e., SAEs requires the same data distribution), or the dataset that was used to train NLLB. However, the `allenai/nllb` dataset available in Huggingface is noisy
Contributor guide
Research direction
No repository file, test, or entry point is identified. Start by confirming which SONAR or NLLB training data can be shared and how it should be packaged on Hugging Face; done means the approved dataset is published with enough provenance and access information to support SAE training.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface, python
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100