asteroid-team / asteroid-team/asteroid

As We Speak: Real-Time Visually Guided Speaker Separation and Localization

Open
#702 0 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
2.6k
Forks
450
PR merge metrics
No merged PRs in 30d

Description

Any plan to implement this kind of paper? For realtime ?

https://ieeexplore.ieee.org/document/9949329

Contributor guide

Open the contributing guide

Research direction

Start by reading the linked IEEE paper and comparing its real-time, visually guided speaker-separation and localization requirements with Asteroid's existing Python and PyTorch capabilities. The issue names no files, tests, or entry points; done would require a decided implementation scope and an agreed way to validate real-time separation and localization.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
audio-video-rtc, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.