huggingface / huggingface/diffusers
[New Pipeline]: Audio-Journey: Visual+LLM-aided Audio Encodec Diffusion
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
### Model/Pipeline/Scheduler description
We efficiently trained an Audio Diffusion model with the aid of Alpaca augmented audio captions using AudioSet labels;
[website](https://audiojourney.github.io/)
[preprint](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_v1.2.pdf)
[Appendix](https://github.com/audiojourney/audiojourney.github.io/blob/main/neurIPS_2023_appendix_v1.3.pdf)
[Implementation](https://github.com/jacksonmichaels/diffusers_with_dataloader)
Weights will be released soon!
### Open source status
- [X] The model implementation is available
- [ ] The model weights are available (Only relevant if addition is not a scheduler).
### Provide useful links for the implementation
@jacksonmichaels
_No response_
Beitragsleitfaden
Rechercherichtung
Start with the linked AudioJourney implementation and read the linked preprint and appendix to understand the proposed pipeline. The issue names no diffusers files or tests; completion would require confirming the model implementation and weights are available and defining how the new pipeline should be integrated.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, pytorch
- Bereich
- audio-video-rtc, machine-learning
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 18/100