flatironinstitute / flatironinstitute/DeepFRI
Questions about data
- Dominant language
- Python
- Stars
- 356
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
First of all, thank you for sharing such a good project.My question may be a little unprofessional.
1.In script data_collection.sh, I have no way to download the file bc-95.out.Whenever I visit this URL, a "Not Found" prompt appears.
After that, the file with the URL is https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-95.txt downloaded on the PDB.
Will the file I downloaded be the same as the one you gave? If it's different,Can you give me other ways to get the right file?
2.The number of mf/bp/cc in the file nrPDB-GO_2019.06.18_annot.tsv is not the same as the number in the jison file of the trained model you gave, is this the case?
Contributor guide
No contributing guide indexed for this repository
Research direction
Inspect data_collection.sh and the referenced bc-95.out URL first, then compare the downloaded clusters-by-entity-95.txt with the expected input. Check nrPDB-GO_2019.06.18_annot.tsv against the trained model's JSON file for the mf/bp/cc counts. Done means documenting whether the sources differ and identifying the correct data source or explaining the count mismatch.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, shell, tensorflow
- Domain
- bioinformatics, data, machine-learning
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100