Benjamin-Lee / Benjamin-Lee/deep-rules
Unsupervised learning for protein sequences
- Dominant language
- HTML
- Stars
- 226
- Forks
- 44
- PR merge metrics
- No merged PRs in 30d
Description
**Have you checked the [list of proposed tips](https://github.com/Benjamin-Lee/deep-rules/issues?q=is%3Aissue+is%3Aopen+label%3Atip) to see if the tip has already been proposed?**
- [x] Yes
**Did you add yourself as a [contributor](https://github.com/Benjamin-Lee/deep-rules/blob/master/contributors.md) by making a pull request if this is your first contribution?**
- [x] Yes, I added myself or am already a contributor
There has been a fair amount of discussion on Twitter the past few days about how to properly evaluate deep learning models that learn representations of protein sequences. This may provide good examples for how to evaluate models. For reference:
- https://twitter.com/larsjuhljensen/status/1124983525873156096
- Comments on https://doi.org/10.1101/622803
I haven't looked at these papers in particular, but it reminds me of related discussions in biochemistry like https://doi.org/10.1021/acs.jcim.7b00403. In that domain, there are pitfalls when dataset splits do not account for chemical similarity.
Contributor guide
Research direction
Start by reviewing the cited Twitter discussion, preprint comments, and biochemistry paper. Determine what guidance the manuscript should add for evaluating representations learned from protein sequences, including dataset splits that account for chemical similarity. Done means the issue's proposed tip is clearly defined and supported by relevant examples.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- bioinformatics, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100