flatironinstitute / flatironinstitute/DeepFRI
Need training dataset for CC functions (16020 available from 27000)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 356
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
Thanks a lot for this amazing work. I just wanted to ask, if the whole dataset for CC functions can be kindly provided. We have 29,902 data for MF and BP. However for CC (PDB-GO ), we could only find 16020 data points from this file (nrPDB-GO_2019.06.18_annot) provided in your pre-processing directory in GitHub.
It will be very helpful to get 29,902 ( i.e. data points) for CC functions as well. Thanks a lot.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in the preprocessing directory with nrPDB-GO_2019.06.18_annot and compare the available CC entries with the requested 29,902 data points. Determine whether the missing records can be provided from the project's source data, and consider the work complete when the full CC training dataset is available or its unavailability is documented.
Written by the indexing model from the issue text.
Assessment
- Domain
- bioinformatics, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100