google-deepmind / google-deepmind/alphafold

Question about the self-distillation dataset

Open
#996 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
14.9k
Forks
2.9k
PR merge metrics
No merged PRs in 30d

Description

I have a question regarding the self-distillation dataset part

![Screenshot from 2024-08-01 11-05-19](https://github.com/user-attachments/assets/9e3fec7e-9b29-4bbc-947a-12460fc07ec5)

- Do you compute an MSA for every sequence of every cluster against the entire Uniclust30 ?

OR

- Do you compute an MSA for the **seed/representative sequence** of every cluster against the entire Uniclust30 ?

and then if a sequence occurs in another (different) cluster's representative MSA you eliminate it ?

thank you for your time and consideration

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.