Whisper encode roughly 4x slower than openai/pytorch
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 4.7k
- Forks
- 536
- Avg merge
- 12h 12m
- Merged PRs (30d)
- 4
Description
Obviously the encoding time is almost a non-issue, only when you are working on very small audio chunks it could even hope to shave off some meaningful total percentage of runtime.
I just wanted to mention it in case it is flying under the radar and there might be a quick fix to it. For example, on RTX 4080 both Linux/Windows the encode takes around 0.08s in ctranslate2 and 0.02s with the openAI reference implementation, same 4x difference on two other systems with RTX 4090 and RTX 4060. Thanks for all the work!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the reported Whisper encoding comparison on the listed RTX systems, using the OpenNMT/CTranslate2 implementation and the OpenAI reference implementation. Measure the roughly 0.08s versus 0.02s encode times and inspect the relevant encoding path to determine whether the difference is actionable; done means the cause is identified and the performance gap is addressed or explained.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100