jcjohnson / jcjohnson/densecap

What is the ground truth when I use natural language queries to retrieve the source image?

Open
#68 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.6k
Forks
424
PR merge metrics
No merged PRs in 30d

Description

In your paper, your dense captioning model can support image retrieval using natural language queries, and can localize these queries in retrieved images. What the ground truth when you calculate R@n?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the paper's description of natural-language image retrieval and its R@n evaluation. Determine which retrieved image or region is treated as ground truth, then document that definition and how the metric is calculated; the payload names no repository file or test to update.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.