facebookresearch / facebookresearch/detectron2
*Interesting* Subset Evaluation Problem (evaluate the provided pretrained model on a subset of COCO, category has different results (AP) even the instances are the same)
- Dominant language
- Python
- Stars
- 34.7k
- Forks
- 7.9k
- PR merge metrics
- No merged PRs in 30d
Description
If you do not know the root cause of the problem, please post according to this template:
## Instructions To Reproduce the Issue:
No code was modified, just change the JSON file for COCO. The mentioned JSON file (instances_val2017_v1.json) for COCO subset can be found here (https://drive.google.com/file/d/1zQOc59t_hX48dSY6UlGAefX7dBAytxlv/view?usp=sharing) which can be used to replace the "instances_val2017.json" under annotation fold of COCO. Using the register_coco_instances() to register this new JSON file will be great. The subset extracts some categories from COCO validation set.
Check https://stackoverflow.com/help/minimal-reproducible-example for how to ask good questions.
Simplify the steps to reproduce the issue using suggestions from the above link, and provide them below:
1. Full runnable code or full changes you made:
```
register_coco_instances("val",
{},
"/mnt/home/jierendeng/coco-manager/instances_val2017_v1.json",
"/mnt/home/jierendeng/datasets/coco/val2017")
```
2. What exact command you run: No Change
3. __Full logs__ or other relevant observations:4
We can find the summary of this subset (518 images, instances_val2017_v1.json) as :
[08/05 08:51:19 d2.data.build]: Distribution of instances among all 80 categories:
| category | #instances | category | #instances | category | #instances |
|:-------------:|:-------------|:------------:|:-------------|:-------------:|:-------------|
| person | 2482 | bicycle | 2 | car | 167 |
| motorcycle | 3 | airplane | 0 | bus | 1 |
| train | 0 | truck | 30 | boat | 21 |
| traffic light | 11 | fire hydrant | 1 | stop sign | 0 |
| parking meter | 1 | bench | 97 | bird | 4 |
| cat | 0 | dog | 32 | horse | 0 |
| sheep | 3 | cow | 0 | elephant | 2 |
| bear | 2 | zebra | 0 | giraffe | 0 |
| backpack | 37 | umbrella | 20 | handbag | 20 |
| tie | 13 | suitcase | 2 | frisbee | 115 |
| skis | 0 | snowboard | 1 | sports ball | 260 |
| kite | 327 | baseball bat | 145 | baseball gl.. | 148 |
| skateboard | 2 | surfboard | 12 | tennis racket | 225 |
| bottle | 54 | wine glass | 5 | cup | 17 |
| fork | 0 | knife | 0 | spoon | 0 |
| bowl | 6 | banana | 0 | apple | 4 |
| sandwich | 1 | orange | 0 | broccoli | 0 |
| carrot | 0 | hot dog | 0 | pizza | 0 |
| donut | 0 | cake | 0 | chair | 327 |
| couch | 3 | potted plant | 9 | bed | 2 |
| dining table | 3 | toilet | 0 | tv | 3 |
| laptop | 3 | mouse | 2 | remote | 1 |
| keyboard | 3 | cell phone | 3 | microwave | 0 |
| oven | 0 | toaster | 0 | sink | 0 |
| refrigerator | 0 | book | 9 | clock | 2 |
| vase | 4 | scissors | 1 | teddy bear | 2 |
| hair drier | 0 | toothbrush | 0 | | |
| total | 4650 | | | | |
[08/05 08:51:41 d2.evaluation.coco_evaluation]: Per-category segm AP:
| category | AP | category | AP | category | AP |
|:--------------|:-------|:-------------|:-------|:---------------|:-------|
| person | 50.783 | bicycle | 3.535 | car | 25.508 |
| motorcycle | 30.297 | airplane | nan | bus | 0.000 |
| train | nan | truck | 22.030 | boat | 20.142 |
| traffic light | 27.867 | fire hydrant | 0.000 | stop sign | nan |
| parking meter | 22.500 | bench | 4.881 | bird | 2.339 |
| cat | nan | dog | 65.184 | horse | nan |
| sheep | 47.228 | cow | nan | elephant | 75.050 |
| bear | 75.248 | zebra | nan | giraffe | nan |
| backpack | 16.116 | umbrella | 17.363 | handbag | 8.117 |
| tie | 16.672 | suitcase | 0.000 | frisbee | 64.202 |
| skis | nan | snowboard | 50.000 | sports ball | 47.925 |
| kite | 32.156 | baseball bat | 25.284 | baseball glove | 39.051 |
| skateboard | 26.733 | surfboard | 12.253 | tennis racket | 53.906 |
| bottle | 27.829 | wine glass | 2.351 | cup | 22.183 |
| fork | nan | knife | nan | spoon | nan |
| bowl | 14.174 | banana | nan | apple | 14.184 |
| sandwich | 0.000 | orange | nan | broccoli | nan |
| carrot | nan | hot dog | nan | pizza | nan |
| donut | nan | cake | nan | chair | 10.913 |
| couch | 55.096 | potted plant | 17.002 | bed | 22.673 |
| dining table | 4.350 | toilet | nan | tv | 49.876 |
| laptop | 86.634 | mouse | 40.396 | remote | 70.000 |
| keyboard | 51.683 | cell phone | 35.380 | microwave | nan |
| oven | nan | toaster | nan | sink | nan |
| refrigerator | nan | book | 29.631 | clock | 0.000 |
| vase | 60.198 | scissors | 70.000 | teddy bear | 48.274 |
| hair drier | nan | toothbrush | nan | | |
The whole validation set (5000 images) has the summary as :
[08/05 08:57:04 d2.data.build]: Distribution of instances among all 80 categories:
| category | #instances | category | #instances | category | #instances |
|:-------------:|:-------------|:------------:|:-------------|:-------------:|:-------------|
| person | 10777 | bicycle | 314 | car | 1918 |
| motorcycle | 367 | airplane | 143 | bus | 283 |
| train | 190 | truck | 414 | boat | 424 |
| traffic light | 634 | fire hydrant | 101 | stop sign | 75 |
| parking meter | 60 | bench | 411 | bird | 427 |
| cat | 202 | dog | 218 | horse | 272 |
| sheep | 354 | cow | 372 | elephant | 252 |
| bear | 71 | zebra | 266 | giraffe | 232 |
| backpack | 371 | umbrella | 407 | handbag | 540 |
| tie | 252 | suitcase | 299 | frisbee | 115 |
| skis | 241 | snowboard | 69 | sports ball | 260 |
| kite | 327 | baseball bat | 145 | baseball gl.. | 148 |
| skateboard | 179 | surfboard | 267 | tennis racket | 225 |
| bottle | 1013 | wine glass | 341 | cup | 895 |
| fork | 215 | knife | 325 | spoon | 253 |
| bowl | 623 | banana | 370 | apple | 236 |
| sandwich | 177 | orange | 285 | broccoli | 312 |
| carrot | 365 | hot dog | 125 | pizza | 284 |
| donut | 328 | cake | 310 | chair | 1771 |
| couch | 261 | potted plant | 342 | bed | 163 |
| dining table | 695 | toilet | 179 | tv | 288 |
| laptop | 231 | mouse | 106 | remote | 283 |
| keyboard | 153 | cell phone | 262 | microwave | 55 |
| oven | 143 | toaster | 9 | sink | 225 |
| refrigerator | 126 | book | 1129 | clock | 267 |
| vase | 274 | scissors | 36 | teddy bear | 190 |
| hair drier | 11 | toothbrush | 57 | | |
| total | 36335 | | | | |
[08/05 08:59:01 d2.evaluation.coco_evaluation]: Per-category segm AP:
| category | AP | category | AP | category | AP |
|:--------------|:-------|:-------------|:-------|:---------------|:-------|
| person | 47.659 | bicycle | 17.969 | car | 41.815 |
| motorcycle | 32.986 | airplane | 49.252 | bus | 63.667 |
| train | 61.038 | truck | 35.089 | boat | 23.022 |
| traffic light | 26.765 | fire hydrant | 62.378 | stop sign | 66.174 |
| parking meter | 45.015 | bench | 17.275 | bird | 30.338 |
| cat | 66.854 | dog | 57.179 | horse | 41.555 |
| sheep | 43.681 | cow | 46.896 | elephant | 55.802 |
| bear | 69.355 | zebra | 56.278 | giraffe | 51.522 |
| backpack | 16.440 | umbrella | 44.650 | handbag | 14.933 |
| tie | 31.189 | suitcase | 39.446 | frisbee | 62.388 |
| skis | 3.222 | snowboard | 21.955 | sports ball | 46.843 |
| kite | 30.874 | baseball bat | 24.689 | baseball glove | 38.559 |
| skateboard | 31.492 | surfboard | 31.171 | tennis racket | 53.682 |
| bottle | 38.238 | wine glass | 30.948 | cup | 41.307 |
| fork | 15.283 | knife | 12.811 | spoon | 12.155 |
| bowl | 40.012 | banana | 19.467 | apple | 20.167 |
| sandwich | 36.739 | orange | 30.311 | broccoli | 21.661 |
| carrot | 18.588 | hot dog | 27.428 | pizza | 50.275 |
| donut | 45.267 | cake | 35.038 | chair | 18.140 |
| couch | 36.135 | potted plant | 22.723 | bed | 32.042 |
| dining table | 16.106 | toilet | 57.366 | tv | 57.325 |
| laptop | 59.190 | mouse | 64.251 | remote | 28.451 |
| keyboard | 50.812 | cell phone | 34.009 | microwave | 55.995 |
| oven | 31.137 | toaster | 43.311 | sink | 35.765 |
| refrigerator | 57.000 | book | 10.240 | clock | 50.323 |
| vase | 36.551 | scissors | 20.594 | teddy bear | 43.648 |
| hair drier | 0.636 | toothbrush | 14.988 | | |
## Expected behavior:
Let's take the "Sports ball" for an example, the subset extract all 260 instances of "Sports ball" from the whole set. If we evaluate the same model(download from model_zoo, COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml) on the two datasets, we expect that the AP should be extractly the same (same model should have the same prediction on the same images). However, the AP is different ( sports ball | 46.843 | for whole set, | sports ball | 47.925 | for the subset ). The difference is significant.
The result is easy to be reproduced with the mentioned JSON file.
## Environment:
irrelevant to environment.
Contributor guide
Research direction
Start by reproducing the reported COCO evaluation with register_coco_instances(), the supplied instances_val2017_v1.json, and the full validation annotations. Compare the per-category instance counts and segmentation AP for the subset and full validation set, especially sports ball, then determine whether the differing results reflect an evaluation issue. Done means the cause is identified and the expected behavior is documented or verified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- computer-vision, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100