huggingface / huggingface/diffusers

[New Pipeline]: Audio-Journey: Visual+LLM-aided Audio Encodec Diffusion

Offen
#3,826 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
community-examples New pipeline/model
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

### Model/Pipeline/Scheduler description

We efficiently trained an Audio Diffusion model with the aid of Alpaca augmented audio captions using AudioSet labels;
[website](https://audiojourney.github.io/)
[preprint](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_v1.2.pdf)
[Appendix](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_appendix_v1.3.pdf)
[Implementation](https://github.com/jacksonmichaels/diffusers_with_dataloader)
Weights will be released soon!

### Open source status

- [X] The model implementation is available
- [ ] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation
@jacksonmichaels

_No response_

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start with the linked AudioJourney implementation and read the linked preprint and appendix to understand the proposed pipeline. The issue names no diffusers files or tests; completion would require confirming the model implementation and weights are available and defining how the new pipeline should be integrated.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python, pytorch
Bereich
audio-video-rtc, machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
18/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.