antirez / antirez/h3.c

RES audio sampling accumulates low-step error on the video sigma grid

Aperta
#6 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C
Stelle
2.7k
Fork
210
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

## Problem

The public `h3_dit_denoise` RES path converts the audio velocity into a
video-sigma denoised estimate and advances it on `sigmas.video`:

```c
audio_denoised = audio_latent +
sigma_video * time_shift_slope(sigma_video) * audio_velocity;

h3_res_step(..., sigmas.video, ...);
```

The audio output is conditioned on the audio timestep, and the repository
already constructs a separate `sigmas.audio` schedule. Expressing the audio
trajectory through the video coordinate is valid in continuous time, but it
introduces substantial finite-step error in sparse RES sampling.

The Euler paths already advance audio directly on `sigmas.audio`.

## Why the discretizations differ

For the current video and audio shifts:

$$
\sigma_a=f(\sigma_v)=\frac{\sigma_v}{4-3\sigma_v}
$$

The existing reparameterized estimate is:

$$
D_v=x+\sigma_v f'(\sigma_v)v_a
$$

The native audio estimate is:

$$
D_a=x+\sigma_a v_a
$$

With constant audio velocity, $D_a$ is constant. $D_v$ remains nonlinear, so
a sparse multistep approximation has more truncation error.

```mermaid
flowchart LR
V["Audio velocity v_a
conditioned on sigma_a"]

V --> O["Video-coordinate estimate
D_v = x + sigma_v f'(sigma_v) v_a"]
O --> OV["RES on sigmas.video"]

V --> N["Audio-coordinate estimate
D_a = x + sigma_a v_a"]
N --> NA["RES on sigmas.audio"]
```

## Reproduction

Use a scalar sample initialized to `1.0` and a constant data-ward velocity of
`0.5`. Integrating from sigma 1 to 0 has the exact endpoint `1.5`.

| Steps | Existing formulation | Native audio formulation | Expected | Existing endpoint error |
|---:|---:|---:|---:|---:|
| 4 | 1.843430758 | 1.500000000 | 1.5 | 22.8954% |
| 7 | 1.620174885 | 1.500000000 | 1.5 | 8.0117% |
| 20 | 1.494925022 | 1.500000238 | 1.5 | 0.3383% |

At four steps, the video-coordinate result has a 22.8954% endpoint error. At
20 steps, the error falls to 0.3383%, making the behavior much less visible at
denser step counts.

These are deterministic solver-state measurements, not decoded-audio quality
measurements.

## Proposed outcome

Pair the velocity-to-denoised conversion with the sigma grid used by RES:

- Video: `video_latent + sigma_video * video_velocity`, then
`sigmas.video`.
- Audio: `audio_latent + sigma_audio * audio_velocity`, then
`sigmas.audio`.

Add a constant-velocity regression at sparse and ordinary step counts. The
Euler paths and schedule construction should remain unchanged.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.