ByteDance-Seed / ByteDance-Seed/Depth-Anything-3
Performance on ScanNet++
- Dominant language
- Python
- Stars
- 6.3k
- Forks
- 702
- PR merge metrics
- No merged PRs in 30d
Description
Dear authors,
Thanks for this great work!
I couldn't obtain good performance when running DepthAnything3 on images from ScanNet++.
For example, for this [four-view sequence](https://polybox.ethz.ch/index.php/s/L8Qa6G28M7MECJw) where I uploaded the processed images (322x504), the online demo cannot predict consistent results, see the misaligned walls.
When I compared the predicted cameras with GT, I found that the prediction has large error in intrinsics. The GT fovX is 124 degree while the predicted fovX is 99 degree.
The performance is surprising as I think DA3 is trained on lots of ScanNet++ dataset or similar indoor scenes.
Thanks so much for your help.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the online-demo result with the linked four-view ScanNet++ sequence and its processed 322x504 images. Compare the predicted cameras with the provided ground truth, focusing on the reported fovX difference of 99 versus 124 degrees and the resulting wall misalignment. Done means identifying and documenting a reproducible cause or confirming the discrepancy with evidence.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- computer-vision, machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 32/100