huggingface / huggingface/diffusers
F5-TTS Integration
- Vorherrschende Sprache
- Python
- Sterne
- 34.5k
- Forks
- 7.3k
- Ø Merge
- 3 T. 3 Std.
- Gemergte PRs (30 T.)
- 91
Beschreibung
### Model/Pipeline/Scheduler description
F5-TTS is a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT).
It has excellent voice cloning capabilities, and audio generation is of quite high quality.
### Open source status
- [X] The model implementation is available.
- [X] The model weights are available (Only relevant if addition is not a scheduler).
### Provide useful links for the implementation
Paper - https://arxiv.org/abs/2410.06885
Code - https://github.com/SWivid/F5-TTS?tab=readme-ov-file
Weights - https://huggingface.co/SWivid/F5-TTS
Author - @SWivid
Beitragsleitfaden
Rechercherichtung
Start by reading the F5-TTS paper and the linked SWivid/F5-TTS implementation, then inspect the model weights on Hugging Face. Compare the existing diffusers architecture for audio generation with F5-TTS to determine the integration points. Done means F5-TTS is supported in diffusers, but the issue does not define specific files, tests, or acceptance criteria.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- huggingface, python, pytorch
- Bereich
- audio-video-rtc, machine-learning
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Aktiv
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 35/100