microsoft / microsoft/foldingdiff

Training my own dataset

Open
#20 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
568
Forks
74
Avg merge
20h 2m
Merged PRs (30d)
1

Description

Hi, thank you for sharing your very interesting folding model!
I want to train a model using my own dataset. Do I just need to put the dataset in the data\cath directory? Also, I would appreciate it if you could explain how to create a compatible Dataset. It seems like PDB files are not compatible.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the data\cath directory and the issue's distinction between PDB files and the expected Dataset format. Document the required dataset structure, how compatibility is determined, and the steps for training on a custom dataset. Done means a newcomer can prepare compatible data and follow the training process without needing an unanswered explanation.

Written by the indexing model from the issue text.

Assessment

Domain
data, documentation, machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.