flatironinstitute / flatironinstitute/DeepFRI

Need training dataset for CC functions (16020 available from 27000)

Open
#53 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
356
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Hi,

Thanks a lot for this amazing work. I just wanted to ask, if the whole dataset for CC functions can be kindly provided. We have 29,902 data for MF and BP. However for CC (PDB-GO ), we could only find 16020 data points from this file (nrPDB-GO_2019.06.18_annot) provided in your pre-processing directory in GitHub.

It will be very helpful to get 29,902 ( i.e. data points) for CC functions as well. Thanks a lot.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in the preprocessing directory with nrPDB-GO_2019.06.18_annot and compare the available CC entries with the requested 29,902 data points. Determine whether the missing records can be provided from the project's source data, and consider the work complete when the full CC training dataset is available or its unavailability is documented.

Written by the indexing model from the issue text.

Assessment

Domain
bioinformatics, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.