huggingface / huggingface/pixparse

[Explore] Where to take vision features?

Open
#8 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
25
Forks
4
PR merge metrics
No merged PRs in 30d

Description

Donut uses the swin v1 features prior to the final LayerNorm layer (model.norm).

For vit right now we are taking features after the final norm, this is usually the case for many downstream applications but not sure what's best here.

We should compare
* final features after norm
* final features without norm
* final features with class token removed (if it's a vit with class token)
* penultimate features (remove one block)
* penultimate feature map (remove one stage, for resolution hierarchical models ie swin, a higher res feat map)
* FPN (for resolution hierarchical models, merge features across resolution via FPN)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.