facebookresearch / facebookresearch/detectron2

*Interesting* Subset Evaluation Problem (evaluate the provided pretrained model on a subset of COCO, category has different results (AP) even the instances are the same)

Open
#4,462 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
34.7k
Forks
7.9k
PR merge metrics
No merged PRs in 30d

Description

If you do not know the root cause of the problem, please post according to this template:

## Instructions To Reproduce the Issue:

No code was modified, just change the JSON file for COCO. The mentioned JSON file (instances_val2017_v1.json) for COCO subset can be found here (https://drive.google.com/file/d/1zQOc59t_hX48dSY6UlGAefX7dBAytxlv/view?usp=sharing) which can be used to replace the "instances_val2017.json" under annotation fold of COCO. Using the register_coco_instances() to register this new JSON file will be great. The subset extracts some categories from COCO validation set.

Check https://stackoverflow.com/help/minimal-reproducible-example for how to ask good questions.
Simplify the steps to reproduce the issue using suggestions from the above link, and provide them below:

1. Full runnable code or full changes you made:
```
register_coco_instances("val",
{},
"/mnt/home/jierendeng/coco-manager/instances_val2017_v1.json",
"/mnt/home/jierendeng/datasets/coco/val2017")

```
2. What exact command you run: No Change
3. __Full logs__ or other relevant observations:4

We can find the summary of this subset (518 images, instances_val2017_v1.json) as :

[08/05 08:51:19 d2.data.build]: Distribution of instances among all 80 categories:
| category | #instances | category | #instances | category | #instances |
|:-------------:|:-------------|:------------:|:-------------|:-------------:|:-------------|
| person | 2482 | bicycle | 2 | car | 167 |
| motorcycle | 3 | airplane | 0 | bus | 1 |
| train | 0 | truck | 30 | boat | 21 |
| traffic light | 11 | fire hydrant | 1 | stop sign | 0 |
| parking meter | 1 | bench | 97 | bird | 4 |
| cat | 0 | dog | 32 | horse | 0 |
| sheep | 3 | cow | 0 | elephant | 2 |
| bear | 2 | zebra | 0 | giraffe | 0 |
| backpack | 37 | umbrella | 20 | handbag | 20 |
| tie | 13 | suitcase | 2 | frisbee | 115 |
| skis | 0 | snowboard | 1 | sports ball | 260 |
| kite | 327 | baseball bat | 145 | baseball gl.. | 148 |
| skateboard | 2 | surfboard | 12 | tennis racket | 225 |
| bottle | 54 | wine glass | 5 | cup | 17 |
| fork | 0 | knife | 0 | spoon | 0 |
| bowl | 6 | banana | 0 | apple | 4 |
| sandwich | 1 | orange | 0 | broccoli | 0 |
| carrot | 0 | hot dog | 0 | pizza | 0 |
| donut | 0 | cake | 0 | chair | 327 |
| couch | 3 | potted plant | 9 | bed | 2 |
| dining table | 3 | toilet | 0 | tv | 3 |
| laptop | 3 | mouse | 2 | remote | 1 |
| keyboard | 3 | cell phone | 3 | microwave | 0 |
| oven | 0 | toaster | 0 | sink | 0 |
| refrigerator | 0 | book | 9 | clock | 2 |
| vase | 4 | scissors | 1 | teddy bear | 2 |
| hair drier | 0 | toothbrush | 0 | | |
| total | 4650 | | | | |
[08/05 08:51:41 d2.evaluation.coco_evaluation]: Per-category segm AP:
| category | AP | category | AP | category | AP |
|:--------------|:-------|:-------------|:-------|:---------------|:-------|
| person | 50.783 | bicycle | 3.535 | car | 25.508 |
| motorcycle | 30.297 | airplane | nan | bus | 0.000 |
| train | nan | truck | 22.030 | boat | 20.142 |
| traffic light | 27.867 | fire hydrant | 0.000 | stop sign | nan |
| parking meter | 22.500 | bench | 4.881 | bird | 2.339 |
| cat | nan | dog | 65.184 | horse | nan |
| sheep | 47.228 | cow | nan | elephant | 75.050 |
| bear | 75.248 | zebra | nan | giraffe | nan |
| backpack | 16.116 | umbrella | 17.363 | handbag | 8.117 |
| tie | 16.672 | suitcase | 0.000 | frisbee | 64.202 |
| skis | nan | snowboard | 50.000 | sports ball | 47.925 |
| kite | 32.156 | baseball bat | 25.284 | baseball glove | 39.051 |
| skateboard | 26.733 | surfboard | 12.253 | tennis racket | 53.906 |
| bottle | 27.829 | wine glass | 2.351 | cup | 22.183 |
| fork | nan | knife | nan | spoon | nan |
| bowl | 14.174 | banana | nan | apple | 14.184 |
| sandwich | 0.000 | orange | nan | broccoli | nan |
| carrot | nan | hot dog | nan | pizza | nan |
| donut | nan | cake | nan | chair | 10.913 |
| couch | 55.096 | potted plant | 17.002 | bed | 22.673 |
| dining table | 4.350 | toilet | nan | tv | 49.876 |
| laptop | 86.634 | mouse | 40.396 | remote | 70.000 |
| keyboard | 51.683 | cell phone | 35.380 | microwave | nan |
| oven | nan | toaster | nan | sink | nan |
| refrigerator | nan | book | 29.631 | clock | 0.000 |
| vase | 60.198 | scissors | 70.000 | teddy bear | 48.274 |
| hair drier | nan | toothbrush | nan | | |
The whole validation set (5000 images) has the summary as :
[08/05 08:57:04 d2.data.build]: Distribution of instances among all 80 categories:
| category | #instances | category | #instances | category | #instances |
|:-------------:|:-------------|:------------:|:-------------|:-------------:|:-------------|
| person | 10777 | bicycle | 314 | car | 1918 |
| motorcycle | 367 | airplane | 143 | bus | 283 |
| train | 190 | truck | 414 | boat | 424 |
| traffic light | 634 | fire hydrant | 101 | stop sign | 75 |
| parking meter | 60 | bench | 411 | bird | 427 |
| cat | 202 | dog | 218 | horse | 272 |
| sheep | 354 | cow | 372 | elephant | 252 |
| bear | 71 | zebra | 266 | giraffe | 232 |
| backpack | 371 | umbrella | 407 | handbag | 540 |
| tie | 252 | suitcase | 299 | frisbee | 115 |
| skis | 241 | snowboard | 69 | sports ball | 260 |
| kite | 327 | baseball bat | 145 | baseball gl.. | 148 |
| skateboard | 179 | surfboard | 267 | tennis racket | 225 |
| bottle | 1013 | wine glass | 341 | cup | 895 |
| fork | 215 | knife | 325 | spoon | 253 |
| bowl | 623 | banana | 370 | apple | 236 |
| sandwich | 177 | orange | 285 | broccoli | 312 |
| carrot | 365 | hot dog | 125 | pizza | 284 |
| donut | 328 | cake | 310 | chair | 1771 |
| couch | 261 | potted plant | 342 | bed | 163 |
| dining table | 695 | toilet | 179 | tv | 288 |
| laptop | 231 | mouse | 106 | remote | 283 |
| keyboard | 153 | cell phone | 262 | microwave | 55 |
| oven | 143 | toaster | 9 | sink | 225 |
| refrigerator | 126 | book | 1129 | clock | 267 |
| vase | 274 | scissors | 36 | teddy bear | 190 |
| hair drier | 11 | toothbrush | 57 | | |
| total | 36335 | | | | |
[08/05 08:59:01 d2.evaluation.coco_evaluation]: Per-category segm AP:
| category | AP | category | AP | category | AP |
|:--------------|:-------|:-------------|:-------|:---------------|:-------|
| person | 47.659 | bicycle | 17.969 | car | 41.815 |
| motorcycle | 32.986 | airplane | 49.252 | bus | 63.667 |
| train | 61.038 | truck | 35.089 | boat | 23.022 |
| traffic light | 26.765 | fire hydrant | 62.378 | stop sign | 66.174 |
| parking meter | 45.015 | bench | 17.275 | bird | 30.338 |
| cat | 66.854 | dog | 57.179 | horse | 41.555 |
| sheep | 43.681 | cow | 46.896 | elephant | 55.802 |
| bear | 69.355 | zebra | 56.278 | giraffe | 51.522 |
| backpack | 16.440 | umbrella | 44.650 | handbag | 14.933 |
| tie | 31.189 | suitcase | 39.446 | frisbee | 62.388 |
| skis | 3.222 | snowboard | 21.955 | sports ball | 46.843 |
| kite | 30.874 | baseball bat | 24.689 | baseball glove | 38.559 |
| skateboard | 31.492 | surfboard | 31.171 | tennis racket | 53.682 |
| bottle | 38.238 | wine glass | 30.948 | cup | 41.307 |
| fork | 15.283 | knife | 12.811 | spoon | 12.155 |
| bowl | 40.012 | banana | 19.467 | apple | 20.167 |
| sandwich | 36.739 | orange | 30.311 | broccoli | 21.661 |
| carrot | 18.588 | hot dog | 27.428 | pizza | 50.275 |
| donut | 45.267 | cake | 35.038 | chair | 18.140 |
| couch | 36.135 | potted plant | 22.723 | bed | 32.042 |
| dining table | 16.106 | toilet | 57.366 | tv | 57.325 |
| laptop | 59.190 | mouse | 64.251 | remote | 28.451 |
| keyboard | 50.812 | cell phone | 34.009 | microwave | 55.995 |
| oven | 31.137 | toaster | 43.311 | sink | 35.765 |
| refrigerator | 57.000 | book | 10.240 | clock | 50.323 |
| vase | 36.551 | scissors | 20.594 | teddy bear | 43.648 |
| hair drier | 0.636 | toothbrush | 14.988 | | |

## Expected behavior:
Let's take the "Sports ball" for an example, the subset extract all 260 instances of "Sports ball" from the whole set. If we evaluate the same model(download from model_zoo, COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml) on the two datasets, we expect that the AP should be extractly the same (same model should have the same prediction on the same images). However, the AP is different ( sports ball | 46.843 | for whole set, | sports ball | 47.925 | for the subset ). The difference is significant.

The result is easy to be reproduced with the mentioned JSON file.

## Environment:
irrelevant to environment.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the reported COCO evaluation with register_coco_instances(), the supplied instances_val2017_v1.json, and the full validation annotations. Compare the per-category instance counts and segmentation AP for the subset and full validation set, especially sports ball, then determine whether the differing results reflect an evaluation issue. Done means the cause is identified and the expected behavior is documented or verified.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.