Decreasing Success Rate in Binder Design Across Repeated, Independent Runs with Identical Parameters

Aperta
#391 4 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
25/100
Tipo di issue
Bug
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
python

Direzione di ricerca

Start by reproducing independent RFdiffusion, ProteinMPNN, and ColabDesign pipeline runs with identical contig maps, hotspots, hyperparameters, and diffuser.T values of 50, 150, and 200. Compare run outputs and execution environments while checking whether state, caches, or temporary files persist between runs. Done means identifying the source of the declining success rate or documenting a reliable isolation and reproducibility procedure.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Dear RFdiffusion Developers,

I am writing to report a perplexing issue I've encountered while working on a binder design project. I have been running the RFdiffusion, ProteinMPNN, and ColabDesign pipeline multiple times with completely identical settings, but I'm observing a consistent and significant decrease in the binder success rate over subsequent runs.

For every run, all conditions, including the contig map, hotspot definitions, and other hyperparameters, were kept exactly the same. I performed several separate, independent runs for different iteration counts (diffuser.T=150, diffuser.T=50, and diffuser.T=200), with each run being a fresh execution of the pipeline. Despite this, I've noticed a clear trend of diminishing returns. For example, when using 150 iterations, my first run successfully produced 5 high-quality binders. However, by the fourth independent run with the exact same 150-iteration setting, zero successful binders were generated. This pattern of a declining success rate across subsequent runs was also observed with 50 and 200 iterations.

This is confusing because my expectation was that independent runs with identical inputs should produce stochastically similar outcomes and success rates over a large number of trials. Instead, I am seeing what appears to be a systematic degradation.

My Questions

  1. Is this a known phenomenon? Has anyone else reported a "performance decay" across repeated, independent executions of the binder design pipeline?

  2. Could there be a hidden state, cache, or temporary file that is not being properly cleared between runs, which might be influencing the outcome of later executions?

  3. Since my workflow involves RFdiffusion, ProteinMPNN, and ColabDesign, is it possible that an interaction between these tools is causing this issue over repeated use?

  4. Are there any recommended best practices for ensuring true independence and reproducibility between runs to avoid this kind of degradation?

This issue is quite puzzling, and any insights or suggestions you could offer would be immensely helpful for my project. Thank you for your time and continued development of this incredible tool.

Lingua principale
Python
Stelle
3.1k
Fork
644
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di RosettaCommons/RFdiffusion

Tutte le issue di RosettaCommons/RFdiffusion

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.