LAION-AI / LAION-AI/Open-Assistant
dataset: BigScience Biomedical Datasets
Open
Nobody has claimed this yet.
data
- Dominant language
- Python
- Stars
- 37.4k
- Forks
- 3.3k
- PR merge metrics
- No merged PRs in 30d
Description
I was looking into adding some datasets from the BigBio repository, and I had some questions before proceeding.
- Galactica was trained on a subset of the BigBio corpus. I'm not sure which model the ML team has decided on, but if it is to be Galactica, should I omit these previously seen datasets?
- Galactica uses special tags for amino acid, DNA, and SMILES sequences. Should I also tag these entities, or would that get in the way if OA decides to go with a different representative scheme?
- BigBio is a collection of individual datasets. Should I keep these datasets as separate submissions to OA or should I bundle them all together?
If these datasets are outside the scope of OA, please let me know.
Thanks.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Review the BigBio repository and the Galactica model context named in the issue, then consult Open Assistant maintainers about dataset overlap, entity tagging, and submission bundling. Done means an agreed scope and submission structure; the issue does not name implementation files or tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100