Project-MONAI / Project-MONAI/MONAI

HausdorffDistanceMetric percentile returns NaN for a missed prediction, dropping failures from the dataset average

Open
#9,095 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.7k
Forks
1.6k
Avg merge
5d 1h
Merged PRs (30d)
20

Description

Describe the bug

HausdorffDistanceMetric(percentile=...) returns nan when one of the two masks is empty, where percentile=None returns inf for the same input.

nan is not a quieter way of saying inf here. It is this metric's "not applicable" sentinel — it is what both-masks-empty returns — and do_metric_reduction excludes it from the average. A prediction that missed the structure entirely is therefore removed from a dataset score rather than counted as the worst case, and the reported HD95 improves as the model finds fewer structures.

To Reproduce

import torch
from monai.metrics import HausdorffDistanceMetric

gt = torch.zeros(1, 1, 64, 64)
gt[0, 0, 20:28, 20:28] = 1.0
pred = torch.zeros(1, 1, 64, 64)          # the model found nothing

for percentile in (None, 50, 95, 99, 100):
    m = HausdorffDistanceMetric(include_background=True, percentile=percentile)
    m(y_pred=pred, y=gt)
    print(percentile, m.get_buffer().flatten().tolist())

# None -> [inf]
# 50   -> [nan]
# 95   -> [nan]
# 99   -> [nan]
# 100  -> [nan]

The effect over a dataset. 100 images, every non-empty prediction identically 4px off, so the only variable is how many images the model misses:

images missed reported HD95 cases averaged (get_not_nans)
0 4.000 100
50 4.000 50
90 4.000 10
99 4.000 1

A model that finds nothing in 99 images out of 100 reports the same HD95 as one that finds the structure every time. A runnable script that prints the whole table is at https://github.com/asifuddin01/A-PR.

Expected behavior

inf, matching percentile=None, so the case is scored rather than discarded. nan stays reserved for both masks empty, where there is genuinely nothing to measure.

Cause

get_surface_distance reports an infinite distance for every boundary voxel when a mask is empty, so an all-infinite tensor reaches _compute_percentile_hausdorff_distance. torch.quantile interpolates linearly between the two order statistics straddling the requested rank:

inf + (inf - inf) * 0.95  ->  nan

The quantile of a constant sequence is that constant, so the nan is an artefact of the interpolation rather than a property of the distances. percentile=None escapes it because .max() does not interpolate.

Environment

MONAI version: 1.6.0
Numpy version: 2.5.2
Pytorch version: 2.14.0
MONAI flags: HAS_EXT = False, USE_COMPILED = False, USE_META_DICT = False
MONAI rev id: eccefc57550b111ed781d82249dfe77872a0e918

Also reproduced on current dev at c0d1ec1.

Additional context

Only the percentile path is affected. SurfaceDistanceMetric returns inf for the same input and SurfaceDiceMetric returns 0.0; both are already honest about a miss.

Separately, #9033 fixes percentile=0 being treated as unset by the if not percentile: guard. It does not cover this: it adds a .min() branch above the torch.quantile call and leaves the interpolation as it is. The two are independent.

I have a fix and regression tests ready and will open a PR against this issue.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at HausdorffDistanceMetric and trace the percentile path through get_surface_distance and _compute_percentile_hausdorff_distance, then check how do_metric_reduction handles the result. Add regression coverage for one empty mask with percentile values and verify the result is inf and remains included in the dataset reduction, while both-empty masks still produce nan.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.