facebookresearch / facebookresearch/detectron2

Semantic Segmentation - Multiclass Classification - Implementation for per Class Accuracy mixed up with per Class Recall?

Open
#4,862 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
34.7k
Forks
7.9k
PR merge metrics
No merged PRs in 30d

Description

### Discussed in https://github.com/facebookresearch/detectron2/discussions/4861

Originally posted by **biggeR-data** March 16, 2023
Hey everyone, I am using Detectron2 with a custom dataset for semantic segmentation. My dataset contains multiple classes so it is not a binary classification problem.

I checked the implementation details for the evaluation metrics in the script [detectron2/evaluation/sem_seg_evaluation.py](https://github.com/facebookresearch/detectron2/blob/main/detectron2/evaluation/sem_seg_evaluation.py). There I stumbled across the per class accuracy implementation which seems to be confused with the per class recall.

Here're the [code parts in question](https://github.com/facebookresearch/detectron2/blob/main/detectron2/evaluation/sem_seg_evaluation.py#L186-L193):
```python
acc = np.full(self._num_classes, np.nan, dtype=np.float)
# ...
tp = self._conf_matrix.diagonal()[:-1].astype(np.float)
pos_gt = np.sum(self._conf_matrix[:-1, :-1], axis=0).astype(np.float)
# ...
acc_valid = pos_gt > 0
acc[acc_valid] = tp[acc_valid] / pos_gt[acc_valid]
```

Judging by the naming of the variables `pos_gt` represents the Ground Truths / Actuals. This means the Actuals are in the columns of the confusion matrix and the Predictions are in the rows of the confusion matrix.
> Note: This notation differs in orientation from the [Wikipedia Definition of a Confusion Matrix](https://en.wikipedia.org/wiki/Confusion_matrix). If you want to calculate Metrics according to the Wikipedia layout you would need to transpose the Confusion Matrix given in the evaluation script.

I will stick to the orientation provided by detectron2 with my following examples to avoid confusion.

Looking at the code the Accuracy per class is calculated by dividing TP by the Actual Positives (named `P` in Wikipedia's entry). This does not correspond to the definition of the Accuracy measure. The Accuracy is defined as:
```
(TP + TN) / n
```
also known by this Formula:
```
(TP + TN) / (TP + FP + TN + FN)
```

The Recall is defined as:
```
TP / (TP + FN) = TP / P
```
which is exactly the formula used to calculate the per class 'accuracy' in [Line 193](https://github.com/facebookresearch/detectron2/blob/main/detectron2/evaluation/sem_seg_evaluation.py#L193).

I have searched for articles covering Multiclass Classification where per class accuracy and per class
recall are covered however the sources for this are rather scarce. I did find a [comment on Stackoverflow](https://stackoverflow.com/questions/39770376/scikit-learn-get-accuracy-scores-for-each-class/65673016#comment118400712_50977153) claiming per class accuracy and per class recall are the same for multiclass classification. On the other hand I found an [example](http://rasbt.github.io/mlxtend/user_guide/evaluate/accuracy_score/) where I continued the given example and arrived at the conclusion that per class accuracy is in fact not the same as per class recall.

Please refer to this screenshot of the continuation of the example: ![Screenshot 2023-03-16 at 12 19 47](https://user-images.githubusercontent.com/75607150/225601537-04480d19-e8e0-417c-902b-dd33c8dd25f9.png)

So I guess my question is:
Is the current implementation for per class accuracy in `sem_seg_evaluation.py` mixed up with per class recall given a multiclass classification problem?

Contributor guide

Open the contributing guide

Research direction

Start by reading detectron2/evaluation/sem_seg_evaluation.py, especially lines 186-193, and review the linked discussion and metric definitions in the issue. Verify the confusion-matrix orientation and whether the reported per-class metric matches its name; done means reaching a documented decision about the expected metric and aligning the implementation if it is incorrect.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.