alibaba / alibaba/LucaProt

Questions on Training LucaProt Model with New Data

Open
#105 18 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
231
Forks
43
PR merge metrics
No merged PRs in 30d

Description

Hi Team,

Thanks for your outstanding work! We are interested in updating your published LucaProt model by incorporating newly available data. Here are some questions we currently have:

1. If we fine-tune starting from the best checkpoint of the published LucaProt model, could using too little new data actually degrade performance? If so, is there a recommended minimum dataset size for effective updating? Do you have any suggestions regarding hyperparameter settings for that size?

2. While we are choosing the new dataset, Are there recommended proportions for different sequence categories (e.g., viral RdRP, bacterial RdRP, eukaryotic RdRP, viral non-RdRP, etc.)? Is there an expected ratio between positive (1) and negative (0) samples?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or training entry point. First clarify the intended data-update procedure and required dataset and hyperparameter guidance with maintainers; completion cannot be defined from the current issue.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.