ByteDance-Seed / ByteDance-Seed/Depth-Anything-3

Trying to match LIDAR measurements

Open
#202 2 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
6.3k
Forks
702
PR merge metrics
No merged PRs in 30d

Description

Hello,

Thank you for your fantastic work!

I am currently using the DA3NESTED-GIANT-LARGE model to infer metric depth. I process videos filmed by iPhone 15 Pro, iPhone 16 Pro, and iPhone 14 Pro Max, all equipped with LIDAR and mounted on tripods. All the videos I use represent static scenes. I process each video at once as series of frames.

As I have the camera intrinsics and extrinsics and depth ground truth, I've noticed the model's predicted depths differ from the ground truth. While the differences are sometimes marginal, in other instances, they vary substantially, up to 3 times. Moreover, the depths of static objects fluctuate between frames. The predicted intrinsics and extrinsics do not match the ground truth either. At the same time, on a **relative scale** depth appear reasonable and accurate. Is there a method to align the predicted depth with the ground truth from LIDAR?

Once again, thank you for your exceptional work.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the DA3NESTED-GIANT-LARGE inference described with the iPhone video frames, camera calibration, and LIDAR ground truth. No repository files or tests are named, so first locate the model inference and depth-alignment entry points. Done would require a confirmed explanation or an agreed method for aligning predictions and handling frame-to-frame variation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.