RES audio sampling accumulates low-step error on the video sigma grid
- Ngôn ngữ chính
- C
- Star
- 2.7k
- Fork
- 210
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Mô tả
## Problem
The public `h3_dit_denoise` RES path converts the audio velocity into a
video-sigma denoised estimate and advances it on `sigmas.video`:
```c
audio_denoised = audio_latent +
sigma_video * time_shift_slope(sigma_video) * audio_velocity;
h3_res_step(..., sigmas.video, ...);
```
The audio output is conditioned on the audio timestep, and the repository
already constructs a separate `sigmas.audio` schedule. Expressing the audio
trajectory through the video coordinate is valid in continuous time, but it
introduces substantial finite-step error in sparse RES sampling.
The Euler paths already advance audio directly on `sigmas.audio`.
## Why the discretizations differ
For the current video and audio shifts:
$$
\sigma_a=f(\sigma_v)=\frac{\sigma_v}{4-3\sigma_v}
$$
The existing reparameterized estimate is:
$$
D_v=x+\sigma_v f'(\sigma_v)v_a
$$
The native audio estimate is:
$$
D_a=x+\sigma_a v_a
$$
With constant audio velocity, $D_a$ is constant. $D_v$ remains nonlinear, so
a sparse multistep approximation has more truncation error.
```mermaid
flowchart LR
V["Audio velocity v_a
conditioned on sigma_a"]
V --> O["Video-coordinate estimate
D_v = x + sigma_v f'(sigma_v) v_a"]
O --> OV["RES on sigmas.video"]
V --> N["Audio-coordinate estimate
D_a = x + sigma_a v_a"]
N --> NA["RES on sigmas.audio"]
```
## Reproduction
Use a scalar sample initialized to `1.0` and a constant data-ward velocity of
`0.5`. Integrating from sigma 1 to 0 has the exact endpoint `1.5`.
| Steps | Existing formulation | Native audio formulation | Expected | Existing endpoint error |
|---:|---:|---:|---:|---:|
| 4 | 1.843430758 | 1.500000000 | 1.5 | 22.8954% |
| 7 | 1.620174885 | 1.500000000 | 1.5 | 8.0117% |
| 20 | 1.494925022 | 1.500000238 | 1.5 | 0.3383% |
At four steps, the video-coordinate result has a 22.8954% endpoint error. At
20 steps, the error falls to 0.3383%, making the behavior much less visible at
denser step counts.
These are deterministic solver-state measurements, not decoded-audio quality
measurements.
## Proposed outcome
Pair the velocity-to-denoised conversion with the sigma grid used by RES:
- Video: `video_latent + sigma_video * video_velocity`, then
`sigmas.video`.
- Audio: `audio_latent + sigma_audio * audio_velocity`, then
`sigmas.audio`.
Add a constant-velocity regression at sparse and ordinary step counts. The
Euler paths and schedule construction should remain unchanged.
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Đánh giá
Issue này chưa được đánh giá.