AlibabaResearch / AlibabaResearch/AdvancedLiterateMachinery

VGT evaluation result not matching for DoclayNet.

Open
#188 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
1.8k
Forks
195
PR merge metrics
No merged PRs in 30d

Description

Hi,

Thank you for creating and sharing the Vision Grid Transformer repository.

I am currently trying to evaluate the model on the DocLayNet test dataset in order to replicate the published results (mAP of 83.7). I am using the weights available here: [doclaynet_VGT_model.pth](https://github.com/AlibabaResearch/AdvancedLiterateMachinery/releases/download/v1.3.0-VGT-release/doclaynet_VGT_model.pth).

I executed the evaluation using the following command:

bash
Copy code
python path/to/train_VGT.py --config-file VGT/object_detection/Configs/cascade/doclaynet_VGT_cascade_PTM.yaml --eval-only --num-gpus 1 MODEL.WEIGHTS VGT/downloads/weights/doclaynet_VGT_model.pth OUTPUT_DIR VGT/AdvancedLiterateMachinery/DocumentUnderstanding/VGT/downloads
However, the results I obtained differ from the published ones. Please see the attached matrix for reference:

![image](https://github.com/user-attachments/assets/9b6a9443-62b2-441f-83f8-d3fb6be02a51)

Could you please advise if I might be overlooking something in the evaluation process?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.