tensorflow / tensorflow/models

Tensorflow Object Detection API for Faster RCNN training too slow

Open
#7,629 0 comments 0 reactions 3 assignees View on GitHub

@pkulzc is already working on this.

Since Jun 21, 2020.

models:research:odapi type:support
Dominant language
Python
Stars
77.7k
Forks
44.8k
PR merge metrics
No merged PRs in 30d

Description

Please go to Stack Overflow for help and support:

http://stackoverflow.com/questions/tagged/tensorflow

Also, please understand that many of the models included in this repository are experimental and research-style code. If you open a GitHub issue, here is our policy:

  1. It must be a bug, a feature request, or a significant problem with documentation (for small docs fixes please send a PR instead).
  2. The form below must be filled out.

Here's why we have that policy: TensorFlow developers respond to issues. We want to focus on work that benefits the whole community, e.g., fixing bugs and adding features. Support only helps individuals. GitHub also notifies thousands of people when issues are filed. We want them to see you communicating an interesting problem, rather than being redirected to Stack Overflow.


System information
  • **What is the top-level directory of the model you are usingresearch/object_detection/:
  • Have I written custom code (as opposed to using a stock example script provided in TensorFlow):
  • **OS Platform and Distribution (e.g., Linux Ubuntu 16.04)Red Hat Enterprise Linux Server release 7.7:
  • **TensorFlow installed from (source or binary)tensorflow/1.12:
  • TensorFlow version (use command below):
  • Bazel version (if compiling from source):
  • **CUDA/cuDNN versioncuda/10.0 cudnn/7.4:
  • **GPU model and memoryNVIDIA Titan XP, NVIDIA TitanX:
  • **Exact command to reproducepython model_main.py --num_train_steps=29000 --model_dir=UAV-training/ --pipeline_config_path=UAV-training/faster_rcnn_inception_v2_UAV.config --logtostderr:
Describe the problem

I started training on Oct 3rd and so far (Oct 8th) the training has progressed too slowly with around 5000 batches that has been completed. The total number of steps I wanted to train it on is 29000. Is it normal behavior? Please help.

Source code / logs

UAV-training/faster_rcnn_inception_v2_UAV.config

Faster R-CNN with Inception v2, configured for UAV Dataset.

Users should configure the fine_tune_checkpoint field in the train config as

well as the label_map_path and input_path fields in the train_input_reader and

eval_input_reader. Search for "PATH_TO_BE_CONFIGURED" to find the fields that

should be configured.

model {
faster_rcnn {
num_classes: 31
image_resizer {
keep_aspect_ratio_resizer {
min_dimension: 600
max_dimension: 1024
}
}
feature_extractor {
type: 'faster_rcnn_inception_v2'
first_stage_features_stride: 16
}
first_stage_anchor_generator {
grid_anchor_generator {
scales: [0.25, 0.5, 1.0, 2.0]
aspect_ratios: [0.5, 1.0, 2.0]
height_stride: 16
width_stride: 16
}
}
first_stage_box_predictor_conv_hyperparams {
op: CONV
regularizer {
l2_regularizer {
weight: 0.0
}
}
initializer {
truncated_normal_initializer {
stddev: 0.01
}
}
}
first_stage_nms_score_threshold: 0.0
first_stage_nms_iou_threshold: 0.7
first_stage_max_proposals: 300
first_stage_localization_loss_weight: 2.0
first_stage_objectness_loss_weight: 1.0
initial_crop_size: 14
maxpool_kernel_size: 2
maxpool_stride: 2
second_stage_box_predictor {
mask_rcnn_box_predictor {
use_dropout: false
dropout_keep_probability: 1.0
fc_hyperparams {
op: FC
regularizer {
l2_regularizer {
weight: 0.0
}
}
initializer {
variance_scaling_initializer {
factor: 1.0
uniform: true
mode: FAN_AVG
}
}
}
}
}
second_stage_post_processing {
batch_non_max_suppression {
score_threshold: 0.0
iou_threshold: 0.6
max_detections_per_class: 100
max_total_detections: 300
}
score_converter: SOFTMAX
}
second_stage_localization_loss_weight: 2.0
second_stage_classification_loss_weight: 1.0
}
}

train_config: {
batch_size: 1
optimizer {
momentum_optimizer: {
learning_rate: {
manual_step_learning_rate {
initial_learning_rate: 0.0002
schedule {
step: 900000
learning_rate: .00002
}
schedule {
step: 1200000
learning_rate: .000002
}
}
}
momentum_optimizer_value: 0.9
}
use_moving_average: false
}
gradient_clipping_by_norm: 10.0
fine_tune_checkpoint: "/afs/crc.nd.edu/user/s/sbanerj2/models/research/object_detection/faster_rcnn_inception_v2_coco_2018_01_28/model.ckpt"
from_detection_checkpoint: true
load_all_detection_checkpoint_vars: true

Note: The below line limits the training process to 200K steps, which we

empirically found to be sufficient enough to train the pets dataset. This

effectively bypasses the learning rate schedule (the learning rate will

never decay). Remove the below line to train indefinitely.

num_steps: 29000
data_augmentation_options {
random_horizontal_flip {
}
}
}

train_input_reader: {
tf_record_input_reader {
input_path: "/afs/crc.nd.edu/group/cvrl/scratch_7/UAV/UG2_Challenge2019/Object_Detection/OBJDETdevkit/UAV-train.record"
}
label_map_path: "/afs/crc.nd.edu/user/s/sbanerj2/models/research/object_detection/data/UAV_label_map.pbtxt"
}

eval_config: {
metrics_set: "coco_detection_metrics"
num_examples: 7237
}

eval_input_reader: {
tf_record_input_reader {
input_path: "/afs/crc.nd.edu/group/cvrl/scratch_7/UAV/UG2_Challenge2019/Object_Detection/OBJDETdevkit/UAV-val.record"
}
label_map_path: "/afs/crc.nd.edu/user/s/sbanerj2/models/research/object_detection/data/UAV_label_map.pbtxt"
shuffle: false
num_readers: 1
}

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.