aws / aws/amazon-sagemaker-examples

Limit Audio File Whisper Large v2

Open
#4,475 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
11k
Forks
7k
Avg merge
8h 29m
Merged PRs (30d)
8

Description

I have deployed the whisper largev2 model through JumpStart with SageMaker. It has deployed an endpoint which works correctly. However, in the endpoint's response, it only transcribes a part of the audio, which corresponds to 30 seconds of the audio. When I have tested the model locally, it works with audios of any length. Is it possible to process longer audios in a single call to the endpoint in some way?

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the SageMaker JumpStart Whisper Large v2 endpoint behavior with audio longer than 30 seconds and compare it with the local model run. Check whether the endpoint request, response, or documented inference behavior defines this limit; done means a documented, reproducible explanation or an explicitly scoped change request.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws
Domain
cloud, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.