facebookresearch / facebookresearch/perception_models

Clarification on Preprocessing for Images of Different Resolutions

Open
#118 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2.4k
Forks
162
PR merge metrics
No merged PRs in 30d

Description

Hi. thanks for open-sourcing the amazing Perception Encoder! Could you clarify two points about image preprocessing, especially referencing Table 33's description ("trained with dynamic tiling for different image sizes and aspect ratio; up to 4 image tiles of the encoder’s native resolution + a thumbnail"):
1. When is the input resized to fixed native sizes (e.g., 336px for L-scale, 448px for G-scale)?
2. When is dynamic tiling applied instead?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.