huggingface / huggingface/pixparse
[BUG] Donut style training (Cruller) w/ vit appears unstable at higher resolution
Open
bug
- Dominant language
- Python
- Stars
- 25
- Forks
- 4
- PR merge metrics
- No merged PRs in 30d
Description
Moving higher resolution with a base vit model there appears to be training issues vs lower resolution. As with lower res, the training appears to converge initially, starting to look good and then there is a sudden loss of ability.
@molbap observed this where the OCR metrics appeared to jump back to 100% error (correct me if wrong)
I also observed behaviour similar to this but in my case my train loss also jumped up to a higher value and did not improve afterwards.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.