google-deepmind / google-deepmind/alphafold
Question about the self-distillation dataset
Open
- Dominant language
- Python
- Stars
- 14.9k
- Forks
- 2.9k
- PR merge metrics
- No merged PRs in 30d
Description
I have a question regarding the self-distillation dataset part

- Do you compute an MSA for every sequence of every cluster against the entire Uniclust30 ?
OR
- Do you compute an MSA for the **seed/representative sequence** of every cluster against the entire Uniclust30 ?
and then if a sequence occurs in another (different) cluster's representative MSA you eliminate it ?
thank you for your time and consideration
Contributor guide
Assessment
This issue has not been assessed yet.