tensorflow / tensorflow/models

Mask loss in ssd_meta_arch

Open
#8,552 2 comments 0 reactions 3 assignees View on GitHub

@pkulzc is already working on this.

Since May 22, 2020.

models:research:odapi type:feature
Dominant language
Python
Stars
77.7k
Forks
44.8k
PR merge metrics
No merged PRs in 30d

Description

I'm adding mask head to the ssd meta architecture. For this I've added necessary changes in the protos and convolution head. Now, the prediction_dict has a mask_predictions in it using ConvolutionalMaskHead with little custom changes. SSD_meta_arch postprocess seems to handle mask head. But when looking at the loss function, it seems that it's not handling the mask loss.

In the mask-rcnn the _loss_box_classifier is defining the mask loss in the following way:

      second_stage_mask_loss = None
      if prediction_masks is not None:
        if groundtruth_masks_list is None:
          raise ValueError('Groundtruth instance masks not provided. '
                           'Please configure input reader.')

        if not self._is_training:
          (proposal_boxes, proposal_boxlists, paddings_indicator,
           one_hot_flat_cls_targets_with_background
          ) = self._get_mask_proposal_boxes_and_classes(
              detection_boxes, num_detections, image_shape,
              groundtruth_boxlists, groundtruth_classes_with_background_list,
              groundtruth_weights_list)
        unmatched_mask_label = tf.zeros(image_shape[1:3], dtype=tf.float32)
        (batch_mask_targets, _, _, batch_mask_target_weights,
         _) = target_assigner.batch_assign_targets(
             target_assigner=self._detector_target_assigner,
             anchors_batch=proposal_boxlists,
             gt_box_batch=groundtruth_boxlists,
             gt_class_targets_batch=groundtruth_masks_list,
             unmatched_class_label=unmatched_mask_label,
             gt_weights_batch=groundtruth_weights_list)

        # Pad the prediction_masks with to add zeros for background class to be
        # consistent with class predictions.
        if prediction_masks.get_shape().as_list()[1] == 1:
          # Class agnostic masks or masks for one-class prediction. Logic for
          # both cases is the same since background predictions are ignored
          # through the batch_mask_target_weights.
          prediction_masks_masked_by_class_targets = prediction_masks
        else:
          prediction_masks_with_background = tf.pad(
              prediction_masks, [[0, 0], [1, 0], [0, 0], [0, 0]])
          prediction_masks_masked_by_class_targets = tf.boolean_mask(
              prediction_masks_with_background,
              tf.greater(one_hot_flat_cls_targets_with_background, 0))

        mask_height = shape_utils.get_dim_as_int(prediction_masks.shape[2])
        mask_width = shape_utils.get_dim_as_int(prediction_masks.shape[3])
        reshaped_prediction_masks = tf.reshape(
            prediction_masks_masked_by_class_targets,
            [batch_size, -1, mask_height * mask_width])

        batch_mask_targets_shape = tf.shape(batch_mask_targets)
        flat_gt_masks = tf.reshape(batch_mask_targets,
                                   [-1, batch_mask_targets_shape[2],
                                    batch_mask_targets_shape[3]])

        # Use normalized proposals to crop mask targets from image masks.
        flat_normalized_proposals = box_list_ops.to_normalized_coordinates(
            box_list.BoxList(tf.reshape(proposal_boxes, [-1, 4])),
            image_shape[1], image_shape[2], check_range=False).get()

        flat_cropped_gt_mask = self._crop_and_resize_fn(
            tf.expand_dims(flat_gt_masks, -1),
            tf.expand_dims(flat_normalized_proposals, axis=1),
            [mask_height, mask_width])
        # Without stopping gradients into cropped groundtruth masks the
        # performance with 100-padded groundtruth masks when batch size > 1 is
        # about 4% worse.
        # TODO(rathodv): Investigate this since we don't expect any variables
        # upstream of flat_cropped_gt_mask.
        flat_cropped_gt_mask = tf.stop_gradient(flat_cropped_gt_mask)

        batch_cropped_gt_mask = tf.reshape(
            flat_cropped_gt_mask,
            [batch_size, -1, mask_height * mask_width])

        mask_losses_weights = (
            batch_mask_target_weights * tf.cast(paddings_indicator,
                                                dtype=tf.float32))
        mask_losses = self._second_stage_mask_loss(
            reshaped_prediction_masks,
            batch_cropped_gt_mask,
            weights=tf.expand_dims(mask_losses_weights, axis=-1),
            losses_mask=losses_mask)
        total_mask_loss = tf.reduce_sum(mask_losses)
        normalizer = tf.maximum(
            tf.reduce_sum(mask_losses_weights * mask_height * mask_width), 1.0)
        second_stage_mask_loss = total_mask_loss / normalizer

      if second_stage_mask_loss is not None:
        mask_loss = tf.multiply(self._second_stage_mask_loss_weight,
                                second_stage_mask_loss, name='mask_loss')
        loss_dict[mask_loss.op.name] = mask_loss

The tensor of SSD is very much different from that of faster rcnn architecture. I need help with the implementation of mask loss in SSD.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.