microsoft / microsoft/foldingdiff
Training my own dataset
Nobody has claimed this yet.
- Dominant language
- Jupyter Notebook
- Stars
- 568
- Forks
- 74
- Avg merge
- 20h 2m
- Merged PRs (30d)
- 1
Description
Hi, thank you for sharing your very interesting folding model!
I want to train a model using my own dataset. Do I just need to put the dataset in the data\cath directory? Also, I would appreciate it if you could explain how to create a compatible Dataset. It seems like PDB files are not compatible.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the data\cath directory and the issue's distinction between PDB files and the expected Dataset format. Document the required dataset structure, how compatibility is determined, and the steps for training on a custom dataset. Done means a newcomer can prepare compatible data and follow the training process without needing an unanswered explanation.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100