Imageomics / Imageomics/hpc-inference

Batch Open-vocabulary Detection with Grounding Models

Open
#18 2 comments 0 reactions 1 assignee View on GitHub

@NetZissou is already working on this.

Since Nov 7, 2025.

enhancement
Dominant language
Python
Stars
0
Forks
0
PR merge metrics
No merged PRs in 30d

Description

About

Add a batch pipeline that takes

  • (a) an image corpus (folder or Parquet of binary images/URIs) and,
  • (b) one or more text labels, and returns detection boxes (with scores + optional masks) for each image/label using an open-vocabulary grounding model such as OWLv2

Objective

  • Support open-vocabulary text prompts

    • Single label
    • Multiple labels
  • Run efficiently on GPU(s) with batch inference

  • Emit results in interoperable formats with stable schema

Example

One Label Detection

- RGB Image
- Text Label: ["Fish"]
Image

Multi-labels Detection

- RGB Image
- Text Label: ["coffee mug", "plate", "spoon"]
Image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.