huggingface / huggingface/diffusers

Add Tortoise TTS as a pipeline

Abierto
#3,891 16 comentarios 0 reacciones 0 asignados Ver en GitHub
Good second issue New pipeline/model
Lenguaje dominante
Python
Estrellas
34.5k
Forks
7.3k
Merge medio
3 d 3 h
PR fusionados (30 d)
91

Descripción

### Model/Pipeline/Scheduler description

TorToise is a multi-voice text-to-speech system, which describes a way to apply recent advances in the image generative domain to speech synthesis. It would be great to have this model in diffusers.
I would love to contribute this.

### Open source status

- [X] The model implementation is available
- [X] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

Paper - https://arxiv.org/pdf/2305.07243.pdf
Github repo - https://github.com/neonbjb/tortoise-tts

@sanchit-gandhi @Vaibhavs10

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Start with the linked Tortoise TTS paper and the neonbjb/tortoise-tts implementation to understand the model and available weights, then compare them with existing Diffusers pipeline integrations. Done means Tortoise TTS is available as a supported pipeline in Diffusers, with its implementation and usage expectations documented.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python, pytorch
Área
audio-video-rtc, machine-learning
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
30/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.