KavrakiLab / KavrakiLab/dinc-ensemble

Benchmarking

Open
#9 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Jupyter Notebook
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

Benchmarking on a few datasets:

1. Dataset from the previous benchmark paper https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729087/
Includes large ligands and tests the incremental approach success.
2. Ensemble datasets.
Would be nice to show performance on 2 more ensemble datasets (estrogen receptor goes separately)
- CDK2 dataset
- 453 crystal structures in the PDB
- representative ensemble of 5 structures (using EnGens)
- 383 unique ligands
- EnGens analysis: https://colab.research.google.com/drive/1tq_vcp6XBakBPUJ_EavyasVyaWVloTn5
- PI3K dataset
- ensemble from the EnGens paper (recreate)
- look into ligands
- ER dataset
- ~400 crystal structures
- ~300 unique ligands
- ensemble 3 structures

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the previous benchmark paper and the linked EnGens analysis for the CDK2 dataset, then identify how the CDK2, PI3K, and ER ensembles and ligands should be assembled. Done means benchmark results are produced for the prior-paper dataset and the requested ensemble datasets, with the ER dataset handled separately.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook
Domain
bioinformatics, data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.