facebookresearch / facebookresearch/SONAR

SONAR training dataset in Huggingface

Open
#75 9 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
910
Forks
103
PR merge metrics
No merged PRs in 30d

Description

Hi,

My goal is to train Sparse Autoencoders (SAEs). If the team can share the training dataset in huggingface, then it would be possible to train SAEs that learn features from the training dataset

The problem:
1. To train SAEs, you need to use the dataset that was used to train SONAR (i.e., SAEs requires the same data distribution), or the dataset that was used to train NLLB. However, the `allenai/nllb` dataset available in Huggingface is noisy

Image

Contributor guide

Open the contributing guide

Research direction

No repository file, test, or entry point is identified. Start by confirming which SONAR or NLLB training data can be shared and how it should be packaged on Hugging Face; done means the approved dataset is published with enough provenance and access information to support SAE training.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface, python
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.