Imageomics / Imageomics/hpc-inference
Batch Open-vocabulary Detection with Grounding Models
Open
@NetZissou is already working on this.
Since Nov 7, 2025.
enhancement
- Dominant language
- Python
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
About
Add a batch pipeline that takes
- (a) an image corpus (folder or Parquet of binary images/URIs) and,
- (b) one or more text labels, and returns detection boxes (with scores + optional masks) for each image/label using an open-vocabulary grounding model such as OWLv2
Objective
-
Support open-vocabulary text prompts
- Single label
- Multiple labels
-
Run efficiently on GPU(s) with batch inference
-
Emit results in interoperable formats with stable schema
Example
One Label Detection
- RGB Image
- Text Label: ["Fish"]
Multi-labels Detection
- RGB Image
- Text Label: ["coffee mug", "plate", "spoon"]
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.