nf-core / nf-core/proteinannotator

Add REBEAN (Read Embedding Based Enzyme Annotator) from "Deciphering enzymatic potential in metagenomic reads through DNA language model"

Open
#35 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Nextflow
Stars
15
Forks
12
Avg merge
1d 20h
Merged PRs (30d)
2

Description

Description of feature

https://www.biorxiv.org/content/10.1101/2024.12.10.627786v2.full

Earth’s microbial world plays a fundamental role in shaping the biosphere, steering global processes, soil rejuvenation, and ecological fortification. An overwhelming majority of microbial entities, however, remain un(der)studied. Metagenomics stands to elucidate this microbial “dark matter” by directly sequencing the microbial community DNA from environmental samples. Yet, our ability to explore metagenomic sequences is mostly limited to establishing their similarity to curated datasets of organisms or genes/proteins. Aside from the difficulties in establishing such similarity, the reference-based approaches, by definition, forgo discovery of any entities sufficiently unlike the reference collection.

Presenting a paradigm shift, language model-based methods, offer promising avenues for reference-free analysis of metagenomic reads. Here, we introduce two language models, a pretrained foundation model REMME, aimed at understanding the DNA context of metagenomic reads, and the fine-tuned REBEAN model for predicting the enzymatic potential encoded within the read-corresponding genes. By emphasizing function recognition over gene identification, REBEAN is able to label molecular functions carried both by previously explored genes and by new (orphan) sequences. Inherently, REBEAN identifies the functionally relevant parts of a gene even though it is not explicitly trained to do so. It thus expands enzymatic annotation of unassembled metagenomic reads from extreme environments while maintaining consistency with available annotation methods. Here, our comprehensive analysis highlights our models’ potential for metagenomic read annotation and unearthing of novel enzymes, thus enriching our understanding of microbial communities.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no files, tests, or entry points. Start by reading the linked REBEAN paper and mapping its requirements onto the repository; done would require a defined implementation and validation plan for adding the model.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
bioinformatics, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.