huggingface / huggingface/pixparse

[BUG] Donut style training (Cruller) w/ vit appears unstable at higher resolution

Open
#9 2 comments 0 reactions 1 assignee Claimed by @molbap View on GitHub
bug
Dominant language
Python
Stars
25
Forks
4
PR merge metrics
No merged PRs in 30d

Description

Moving higher resolution with a base vit model there appears to be training issues vs lower resolution. As with lower res, the training appears to converge initially, starting to look good and then there is a sudden loss of ability.

@molbap observed this where the OCR metrics appeared to jump back to 100% error (correct me if wrong)

I also observed behaviour similar to this but in my case my train loss also jumped up to a higher value and did not improve afterwards.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.