ByteDance-Seed / ByteDance-Seed/Depth-Anything-3

How the camera poses reported in the tables are obtained?

Open
#42 0 comments 7 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
6.3k
Forks
702
PR merge metrics
No merged PRs in 30d

Description

Hi, thanks for the great work!

Image

In the ablation study table, I noticed that the camera pose accuracy drops significantlywithout dual-DPT.
Could you clarify how the camera poses reported in the tables are obtained?
Are they estimated from the ray maps, or directly predicted by the camera head?
If they are predicted directly by the camera-pose head, could you explain why removing Dual-DPT leads to such a large performance drop?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.