RosettaCommons / RosettaCommons/foundry
Timeline for training data release
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 966
- Forks
- 181
- Avg merge
- 4d 4h
- Merged PRs (30d)
- 2
Description
Hello,
Thank you for open-sourcing this amazing work, AtomWorks will be great for the community. To this end, I was curious if/when the following datasets will be released?
We also develop two new nucleic acid distillation datasets, described within the supplementary methods: a protein-nucleic acid complex distillation set and an RNA distillation set (with 27K examples and 10K examples, respectively)
If not, could more details (Ideally atleast the raw sequences) about how this dataset was created be made available?
Best,
Talal
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the quoted supplementary-methods description in the issue and determine whether the protein–nucleic acid complex and RNA distillation datasets are planned for release. Done means documenting the release status and, if they will not be released, providing the requested dataset-generation details or raw sequences.
Written by the indexing model from the issue text.
Assessment
- Domain
- bioinformatics, data
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100