huggingface / huggingface/diffusers

[New Pipeline]: Audio-Journey: Visual+LLM-aided Audio Encodec Diffusion

Aperta
#3,826 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub
community-examples New pipeline/model
Lingua principale
Python
Stelle
34.5k
Fork
7.3k
Merge medio
3g 3h
PR unite (30g)
91

Descrizione

### Model/Pipeline/Scheduler description

We efficiently trained an Audio Diffusion model with the aid of Alpaca augmented audio captions using AudioSet labels;
[website](https://audiojourney.github.io/)
[preprint](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_v1.2.pdf)
[Appendix](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_appendix_v1.3.pdf)
[Implementation](https://github.com/jacksonmichaels/diffusers_with_dataloader)
Weights will be released soon!

### Open source status

- [X] The model implementation is available
- [ ] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation
@jacksonmichaels

_No response_

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Start with the linked AudioJourney implementation and read the linked preprint and appendix to understand the proposed pipeline. The issue names no diffusers files or tests; completion would require confirming the model implementation and weights are available and defining how the new pipeline should be integrated.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python, pytorch
Ambito
audio-video-rtc, machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
18/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.