LAION-AI / LAION-AI/Open-Assistant

dataset: BigScience Biomedical Datasets

Open
#1,944 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

data
Dominant language
Python
Stars
37.4k
Forks
3.3k
PR merge metrics
No merged PRs in 30d

Description

I was looking into adding some datasets from the BigBio repository, and I had some questions before proceeding.

  1. Galactica was trained on a subset of the BigBio corpus. I'm not sure which model the ML team has decided on, but if it is to be Galactica, should I omit these previously seen datasets?
  2. Galactica uses special tags for amino acid, DNA, and SMILES sequences. Should I also tag these entities, or would that get in the way if OA decides to go with a different representative scheme?
  3. BigBio is a collection of individual datasets. Should I keep these datasets as separate submissions to OA or should I bundle them all together?

If these datasets are outside the scope of OA, please let me know.

Thanks.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Review the BigBio repository and the Galactica model context named in the issue, then consult Open Assistant maintainers about dataset overlap, entity tagging, and submission bundling. Done means an agreed scope and submission structure; the issue does not name implementation files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
huggingface
Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.