huggingface / huggingface/diffusers

F5-TTS Integration

Offen
#10,043 12 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
contributions-welcome help wanted
Vorherrschende Sprache
Python
Sterne
34.5k
Forks
7.3k
Ø Merge
3 T. 3 Std.
Gemergte PRs (30 T.)
91

Beschreibung

### Model/Pipeline/Scheduler description

F5-TTS is a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT).
It has excellent voice cloning capabilities, and audio generation is of quite high quality.

### Open source status

- [X] The model implementation is available.
- [X] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

Paper - https://arxiv.org/abs/2410.06885
Code - https://github.com/SWivid/F5-TTS?tab=readme-ov-file
Weights - https://huggingface.co/SWivid/F5-TTS

Author - @SWivid

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by reading the F5-TTS paper and the linked SWivid/F5-TTS implementation, then inspect the model weights on Hugging Face. Compare the existing diffusers architecture for audio generation with F5-TTS to determine the integration points. Done means F5-TTS is supported in diffusers, but the issue does not define specific files, tests, or acceptance criteria.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
huggingface, python, pytorch
Bereich
audio-video-rtc, machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Aktiv
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.