ByteDance-Seed / ByteDance-Seed/Depth-Anything-3
Trying to match LIDAR measurements
- Dominant language
- Python
- Stars
- 6.3k
- Forks
- 702
- PR merge metrics
- No merged PRs in 30d
Description
Hello,
Thank you for your fantastic work!
I am currently using the DA3NESTED-GIANT-LARGE model to infer metric depth. I process videos filmed by iPhone 15 Pro, iPhone 16 Pro, and iPhone 14 Pro Max, all equipped with LIDAR and mounted on tripods. All the videos I use represent static scenes. I process each video at once as series of frames.
As I have the camera intrinsics and extrinsics and depth ground truth, I've noticed the model's predicted depths differ from the ground truth. While the differences are sometimes marginal, in other instances, they vary substantially, up to 3 times. Moreover, the depths of static objects fluctuate between frames. The predicted intrinsics and extrinsics do not match the ground truth either. At the same time, on a **relative scale** depth appear reasonable and accurate. Is there a method to align the predicted depth with the ground truth from LIDAR?
Once again, thank you for your exceptional work.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the DA3NESTED-GIANT-LARGE inference described with the iPhone video frames, camera calibration, and LIDAR ground truth. No repository files or tests are named, so first locate the model inference and depth-alignment entry points. Done would require a confirmed explanation or an agreed method for aligning predictions and handling frame-to-frame variation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- computer-vision, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100