deepglint / deepglint/MVT

Some problems when using RICE-ViT

Open
#8 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
HTML
Stars
72
Forks
2
PR merge metrics
No merged PRs in 30d

Description

Hello @anxiangsir 🤗

I'm Yiming and study in NTU. Recently I’ve been working with RICE-ViT and trying to reproduce [baseline ](https://cdn-uploads.huggingface.co/production/uploads/6478679d7b370854241b2ad8/Yy09pOusaZ47LofJ27xox.jpeg) built on Qwen2.5-7B-Instruct. I ran into a couple of questions and would really appreciate your help:

### About reproducing ViT-L-14-336px results
I used [rice-vit-large-patch14-560](https://huggingface.co/DeepGlint-AI/rice-vit-large-patch14-560) and modify the `crop_size` and `shortest_edge` in [preprocessor_config](https://huggingface.co/DeepGlint-AI/rice-vit-large-patch14-560/blob/main/preprocessor_config.json) to 336, attempting to match the ViT-L-14-336px setup. Is this the correct way to reproduce the 336px version? If not, where can I find the checkpoint specifically trained for ViT-L-14-336px?

### Which MLCDVisionModel to use
I noticed that there are two version of MLCDVisionModel,

- one in [LLaVA-NEXT](https://github.com/LLaVA-VL/LLaVA-NeXT/blob/main/llava/model/multimodal_encoder/mlcd/vit_rope2d_hf.py#L275),

- another in [transformers](https://github.com/huggingface/transformers/blob/main/src/transformers/models/mlcd/modeling_mlcd.py#L53)

For RICE-ViT, I used the version from Transformers.
Is this the correct choice?

Thanks a lot for your time! 🙏

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.