facebookresearch / facebookresearch/segment-anything

Multimodal input substitution for RGB

Open
#680 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
54.9k
Forks
6.4k
PR merge metrics
No merged PRs in 30d

Description

Hi, I was wondering if it would be feasible to substitute the traditional RGB image style for three different medical modalities during SAM training and inference. For example, instead of having red, green, and blue channels, have three slices of data, each representing a different medical imaging modality (each is 224x224).

Would SAM reasonably be able to learn how to use information from each of these slices without too much fine tuning?

Contributor guide

Open the contributing guide

Research direction

No file, test, or entry point is named. First inspect the repository’s training and inference entry points to locate the RGB-channel assumptions; done would require a concrete implementation scope and validation plan for three medical modalities.

Written by the indexing model from the issue text.

Assessment

Domain
computer-vision, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.