huggingface / huggingface/speechbox

Anyone have demo source code to process file with whisper large model and get outputs as vtt srt?

Open
#7 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
357
Forks
40
PR merge metrics
No merged PRs in 30d

Description

Dear @patrickvonplaten thank you very much for this repo
I am using almost every day to generate subtitles for my videos by using Whisper

However, I need its produced vtt file (basically subtitle output format of transcription)

Currently to fix and improve punctuation, I am using **fullstop-punctuation-multilang-large** (https://huggingface.co/oliverguhr/fullstop-punctuation-multilang-large) but I can't say it is the best

I would like to test your repo however I need demo for full vtt export

**Could you release a demo source code that can output vtt file ? it can have both fixed and raw output of whipser for comparison**

Moreover, I have added you from linkedin if you accept i appreciate : https://www.linkedin.com/in/furkangozukara/

One final thing. I am also very interested in stable diffusion and preparing tutorial videos. I hope that you consider adding my tutorial videos to readme here : https://huggingface.co/runwayml/stable-diffusion-v1-5

And perhaps open back this topic so people can learn? thank you : https://huggingface.co/runwayml/stable-diffusion-v1-5/discussions/66

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.