flatironinstitute / flatironinstitute/DeepFRI

Questions about data

Open
#37 0 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
356
Forks
89
PR merge metrics
No merged PRs in 30d

Description

First of all, thank you for sharing such a good project.My question may be a little unprofessional.
1.In script data_collection.sh, I have no way to download the file bc-95.out.Whenever I visit this URL, a "Not Found" prompt appears.
After that, the file with the URL is https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-95.txt downloaded on the PDB.
Will the file I downloaded be the same as the one you gave? If it's different,Can you give me other ways to get the right file?
2.The number of mf/bp/cc in the file nrPDB-GO_2019.06.18_annot.tsv is not the same as the number in the jison file of the trained model you gave, is this the case?

Contributor guide

No contributing guide indexed for this repository

Research direction

Inspect data_collection.sh and the referenced bc-95.out URL first, then compare the downloaded clusters-by-entity-95.txt with the expected input. Check nrPDB-GO_2019.06.18_annot.tsv against the trained model's JSON file for the mf/bp/cc counts. Done means documenting whether the sources differ and identifying the correct data source or explaining the count mismatch.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, shell, tensorflow
Domain
bioinformatics, data, machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.