aws-samples / aws-samples/amazon-nova-samples
When will Amazon Nova Sonic support the half-cascased architecture and accept a custom TTS for the voice part
Open
- Dominant language
- Jupyter Notebook
- Stars
- 452
- Forks
- 266
- PR merge metrics
- No merged PRs in 30d
Description
I'm using Amazon Nova Sonic 2 to build a Voice AI Agent using Livekit. I would like to use the half-cascaded architecture and use a custom TTS (from Eleven Labs or Cartesia) instead of using the available voices from Amazon.
Popular speech-to-speech models like OpenAI & Gemini Live already support the half-cascaded architecture. I would like Amazon Nova Sonic to also supporting this as that's important for the use case I'm working on.
Is this something that's on the roadmap for Amazon Nova Sonic??
Contributor guide
Assessment
This issue has not been assessed yet.