Questions on Training LucaProt Model with New Data
- Dominant language
- Python
- Stars
- 231
- Forks
- 43
- PR merge metrics
- No merged PRs in 30d
Description
Hi Team,
Thanks for your outstanding work! We are interested in updating your published LucaProt model by incorporating newly available data. Here are some questions we currently have:
1. If we fine-tune starting from the best checkpoint of the published LucaProt model, could using too little new data actually degrade performance? If so, is there a recommended minimum dataset size for effective updating? Do you have any suggestions regarding hyperparameter settings for that size?
2. While we are choosing the new dataset, Are there recommended proportions for different sequence categories (e.g., viral RdRP, bacterial RdRP, eukaryotic RdRP, viral non-RdRP, etc.)? Is there an expected ratio between positive (1) and negative (0) samples?
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or training entry point. First clarify the intended data-update procedure and required dataset and hyperparameter guidance with maintainers; completion cannot be defined from the current issue.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- machine-learning
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100