tensorflow / tensorflow/models
[object_detection]global_step is not available in checkpoint
@pkulzc is already working on this.
Since May 29, 2020.
- Dominant language
- Python
- Stars
- 77.7k
- Forks
- 44.8k
- PR merge metrics
- No merged PRs in 30d
Description
System information
What is the top-level directory of the model you are using: models-master/research/object_detection/
Have I written custom code (as opposed to using a stock example script provided in TensorFlow): Yes
OS Platform and Distribution (e.g., Linux Ubuntu 16.04): Windows10
TensorFlow installed from (source or binary): binary
TensorFlow version (use command below): 1.12.0
Bazel version (if compiling from source): No
CUDA/cuDNN version: No
GPU model and memory: No GPU/16G memory
Exact command to reproduce: No
Problem description:
When I use this command to train:
python model_main.py --pipeline_config_path=training/ssd_inception_v2_coco.config --model_dir=training/ --num_train_steps=200000 --alsologtostderr
The error message is:
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:523: FutureWarning:
Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
_np_qint8 = np.dtype([("qint8", np.int8, 1)])
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:524: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
_np_quint8 = np.dtype([("quint8", np.uint8, 1)])
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:525: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
_np_qint16 = np.dtype([("qint16", np.int16, 1)])
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:526: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
_np_quint16 = np.dtype([("quint16", np.uint16, 1)])
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:527: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
_np_qint32 = np.dtype([("qint32", np.int32, 1)])
E:\Anaconda3\lib\site-packages\tensorflow\python\framework\dtypes.py:532: FutureWarning: Passing (type, 1) or '1type' as a synonym of type is deprecated; in a future version of numpy, it will be understood as (type, (1,)) / '(1,)type'.
np_resource = np.dtype([("resource", np.ubyte, 1)])
E:\test_opencv\models-master\research\object_detection\utils\visualization_utils.py:29: UserWarning:
This call to matplotlib.use() has no effect because the backend has already
been chosen; matplotlib.use() must be called before pylab, matplotlib.pyplot,
or matplotlib.backends is imported for the first time.The backend was originally set to 'Qt5Agg' by the following code:
File "model_main.py", line 26, in
from object_detection import model_lib
File "E:\test_opencv\models-master\research\object_detection\model_lib.py", line 27, in
from object_detection import eval_util
File "E:\test_opencv\models-master\research\object_detection\eval_util.py", line 33, in
from object_detection.metrics import coco_evaluation
File "E:\test_opencv\models-master\research\object_detection\metrics\coco_evaluation.py", line 25, in
from object_detection.metrics import coco_tools
File "E:\test_opencv\models-master\research\object_detection\metrics\coco_tools.py", line 51, in
from pycocotools import coco
File "E:\Anaconda3\lib\site-packages\pycocotools\coco.py", line 49, in
import matplotlib.pyplot as plt
File "E:\Anaconda3\lib\site-packages\matplotlib\pyplot.py", line 69, in
from matplotlib.backends import pylab_setup
File "E:\Anaconda3\lib\site-packages\matplotlib\backends_init_.py", line 14, in
line for line in traceback.format_stack()import matplotlib; matplotlib.use('Agg') # pylint: disable=multiple-statements
WARNING:tensorflow:Forced number of epochs for all eval validations to be 1.
WARNING:tensorflow:Expected number of evaluation epochs is 1, but instead encounteredeval_on_train_input_config.num_epochs= 0. Overwritingnum_epochsto 1.
WARNING:tensorflow:Estimator's model_fn (<function create_model_fn..model_fn at 0x000001C5EB8610D0>) includes params argument, but params are not passed to Estimator.
WARNING:tensorflow:num_readers has been reduced to 1 to match input file shards.
WARNING:tensorflow:From E:\test_opencv\models-master\research\object_detection\builders\dataset_builder.py:86: parallel_interleave (from tensorflow.contrib.data.python.ops.interleave_ops) is deprecated and will be removed in a future version.
Instructions for updating:
Usetf.data.experimental.parallel_interleave(...).
WARNING:tensorflow:From E:\Anaconda3\lib\site-packages\tensorflow\python\ops\sparse_ops.py:1165: sparse_to_dense (from tensorflow.python.ops.sparse_ops) is deprecated and will be removed in a future version.
Instructions for updating:
Create atf.sparse.SparseTensorand usetf.sparse.to_denseinstead.
WARNING:tensorflow:From E:\test_opencv\models-master\research\object_detection\builders\dataset_builder.py:158: batch_and_drop_remainder (from tensorflow.contrib.data.python.ops.batching) is deprecated and will be removed in a future version.
Instructions for updating:
Usetf.data.Dataset.batch(..., drop_remainder=True).
WARNING:root:Variable [BoxPredictor_0/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[273]], model variable shape: [[21]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_0/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 576, 273]], model variable shape: [[3, 3, 576, 21]]. This variable will not be ini
tialized from the checkpoint.
WARNING:root:Variable [BoxPredictor_1/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[546]], model variable shape: [[42]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_1/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 1024, 546]], model variable shape: [[3, 3, 1024, 42]]. This variable will not be i
nitialized from the checkpoint.
WARNING:root:Variable [BoxPredictor_2/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[546]], model variable shape: [[42]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_2/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 512, 546]], model variable shape: [[3, 3, 512, 42]]. This variable will not be ini
tialized from the checkpoint.
WARNING:root:Variable [BoxPredictor_3/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[546]], model variable shape: [[42]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_3/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 256, 546]], model variable shape: [[3, 3, 256, 42]]. This variable will not be ini
tialized from the checkpoint.
WARNING:root:Variable [BoxPredictor_4/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[546]], model variable shape: [[42]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_4/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 256, 546]], model variable shape: [[3, 3, 256, 42]]. This variable will not be ini
tialized from the checkpoint.
WARNING:root:Variable [BoxPredictor_5/ClassPredictor/biases] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[546]], model variable shape: [[42]]. This variable will not be initialized from the check
point.
WARNING:root:Variable [BoxPredictor_5/ClassPredictor/weights] is available in checkpoint, but has an incompatible shape with model variable. Checkpoint shape: [[3, 3, 128, 546]], model variable shape: [[3, 3, 128, 42]]. This variable will not be ini
tialized from the checkpoint.
WARNING:root:Variable [global_step] is not available in checkpoint
The content in ssd_inception_v2_coco.config likes this:
model {
ssd {
num_classes: 6
image_resizer {
fixed_shape_resizer {
height: 300
width: 300
}
}
feature_extractor {
type: "ssd_inception_v2"
depth_multiplier: 1.0
min_depth: 16
conv_hyperparams {
regularizer {
l2_regularizer {
weight: 3.99999989895e-05
}
}
initializer {
truncated_normal_initializer {
mean: 0.0
stddev: 0.0299999993294
}
}
activation: RELU_6
batch_norm {
decay: 0.999700009823
center: true
scale: true
epsilon: 0.0010000000475
train: true
}
}
override_base_feature_extractor_hyperparams: true
}
box_coder {
faster_rcnn_box_coder {
y_scale: 10.0
x_scale: 10.0
height_scale: 5.0
width_scale: 5.0
}
}
matcher {
argmax_matcher {
matched_threshold: 0.5
unmatched_threshold: 0.5
ignore_thresholds: false
negatives_lower_than_unmatched: true
force_match_for_each_row: true
}
}
similarity_calculator {
iou_similarity {
}
}
box_predictor {
convolutional_box_predictor {
conv_hyperparams {
regularizer {
l2_regularizer {
weight: 3.99999989895e-05
}
}
initializer {
truncated_normal_initializer {
mean: 0.0
stddev: 0.0299999993294
}
}
activation: RELU_6
}
min_depth: 0
max_depth: 0
num_layers_before_predictor: 0
use_dropout: false
dropout_keep_probability: 0.800000011921
kernel_size: 3
box_code_size: 4
apply_sigmoid_to_scores: false
}
}
anchor_generator {
ssd_anchor_generator {
num_layers: 6
min_scale: 0.20000000298
max_scale: 0.949999988079
aspect_ratios: 1.0
aspect_ratios: 2.0
aspect_ratios: 0.5
aspect_ratios: 3.0
aspect_ratios: 0.333299994469
reduce_boxes_in_lowest_layer: true
}
}
post_processing {
batch_non_max_suppression {
score_threshold: 0.300000011921
iou_threshold: 0.600000023842
max_detections_per_class: 100
max_total_detections: 100
}
score_converter: SIGMOID
}
normalize_loss_by_num_matches: true
loss {
localization_loss {
weighted_smooth_l1 {
}
}
classification_loss {
weighted_sigmoid {
}
}
hard_example_miner {
num_hard_examples: 3000
iou_threshold: 0.990000009537
loss_type: CLASSIFICATION
max_negatives_per_positive: 3
min_negatives_per_image: 0
}
classification_weight: 1.0
localization_weight: 1.0
}
}
}
train_config {
batch_size: 10
data_augmentation_options {
random_horizontal_flip {
}
}
data_augmentation_options {
ssd_random_crop {
}
}
optimizer {
rms_prop_optimizer {
learning_rate {
exponential_decay_learning_rate {
initial_learning_rate: 0.004
decay_steps: 800720
decay_factor: 0.95
}
}
momentum_optimizer_value: 0.899999976158
decay: 0.899999976158
epsilon: 1.0
}
}
fine_tune_checkpoint: "ssd_inception_v2_coco_2018_01_28/model.ckpt"
fine_tune_checkpoint_type: "detection"
#from_detection_checkpoint: true
load_all_detection_checkpoint_vars: true
num_steps: 200000
}
train_input_reader {
label_map_path: "data/car_bicycle_person_motorcycle_bus_truck.pbtxt"
tf_record_input_reader {
input_path: "data/train.record"
}
}
eval_config {
num_examples: 8000
max_evals: 10
use_moving_averages: false
}
eval_input_reader {
label_map_path: "data/car_bicycle_person_motorcycle_bus_truck.pbtxt"
shuffle: false
num_readers: 1
tf_record_input_reader {
input_path: "data/test.record"
}
}
But after a while the model started training, it is strange.
How to fix it?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.