facebookresearch / facebookresearch/sam-audio

how to adapt this model to separate different speakers

Open
#5 15 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.6k
Forks
330
PR merge metrics
No merged PRs in 30d

Description

Thanks for sharing the amazing work!

How to adapt this model to separate different speakers? e.g., five people in one audio. They are talking one-by-one instead of speaking together.

Contributor guide

Open the contributing guide

Research direction

Start with the repository's inference code and example notebooks, then check how the model handles a single audio input containing multiple speakers. Done would require a documented, reproducible path for separating five sequential speakers, or a clear explanation of whether the current model supports that use case.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
audio-video-rtc, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.