ByteDance-Seed / ByteDance-Seed/Depth-Anything-3
How the camera poses reported in the tables are obtained?
Open
- Dominant language
- Python
- Stars
- 6.3k
- Forks
- 702
- PR merge metrics
- No merged PRs in 30d
Description
Hi, thanks for the great work!
In the ablation study table, I noticed that the camera pose accuracy drops significantlywithout dual-DPT.
Could you clarify how the camera poses reported in the tables are obtained?
Are they estimated from the ray maps, or directly predicted by the camera head?
If they are predicted directly by the camera-pose head, could you explain why removing Dual-DPT leads to such a large performance drop?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.