deepmodeling / deepmodeling/Uni-Mol

[BUG] _Replace With Suitable Title_

Open
#316 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Python
Stars
1.2k
Forks
181
PR merge metrics
No merged PRs in 30d

Description

### Describe the bug

I create a predicting task to finetune 84M unimol2 model with or without checkpoint published. The dataset contains 100M samples from Molecule3D dataset with HOMO label. Howerver the train loss cannot converge correctly. When without checkpoint, the train loss gradually decreasing to 0.1 and suddenly increasing to 0.55 and cannot decrease to 0.1 any more. And the train loss with checkpoint can also jump to 0.55. I've tried a variety of training parametes, but got similar loss curves. Here is two examples using the finetune parameters suggested in README. The task_name is molecule3d_homo and I've allready written it's mean and std in unimol_finetune task_metainfo.

With checkpoint

![Image](https://github.com/user-attachments/assets/918ec8d4-4483-4dcd-a49d-c9077959f635)

Without checkpoint

![Image](https://github.com/user-attachments/assets/1caf6fae-7b6c-442f-b1cc-7bed922350aa)

I wonder what can I do to keep the finetune trainning steady? Thanks a lot.

### Uni-Mol Version

Uni-Mol2

### Expected behavior

The train loss can decrease and keep steady in unimol finetune task.

### To Reproduce

_No response_

### Environment

V100+python 3.9+pytorch 2.0.0

### Additional Context

_No response_

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.