cap F1 dataset and prompts
Ouverte
- Langage dominant
- Python
- Étoiles
- 933
- Forks
- 96
- Métriques de merge des PR
- Aucune PR mergée en 30 j
Description
Hi,
I am trying to reproduce the cap F1 procedure described in the first paragraph of section C in the Molmo and Pixmo paper.
Is it possible to release the dataset used for that evaluation (the 1500 image and their transcripts at least), as well as the prompts used with GPT-4o to compute both precision and recall (enumerating atomic statements, matching them, and checking consistency) ?
Thank you!
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Évaluation
Cette issue n'a pas encore été évaluée.