tensorflow / tensorflow/models

Apply landmark to MoViNet

Open
#13,569 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

type:support
Dominant language
Python
Stars
77.7k
Forks
44.8k
PR merge metrics
No merged PRs in 30d

Description

def build_classifier(batch_size, num_frames, backbone, resolution, num_classes):
# --- Video input ---
video_input = layers.Input(shape=(num_frames, resolution, resolution, 3), batch_size=batch_size, name='video_input')

# Feature extraction from MoViNet backbone
def extract_video_features(x):
    endpoints = backbone(x, training=False)
    x = endpoints['head'] 
    x = tf.squeeze(x, axis=[2, 3]) 
    x = tf.keras.layers.GlobalAveragePooling1D()(x)
    return x

video_features = layers.Lambda(
    extract_video_features,
    output_shape=(480,),
    name="video_features"
)(video_input)

# --- Landmark input ---
landmark_input = layers.Input(shape=(num_frames, 234), batch_size=batch_size, name='landmark_input')
landmark_features = layers.Bidirectional(layers.LSTM(128, return_sequences=False))(landmark_input)
landmark_features = layers.Dense(128, activation='relu')(landmark_features)

# --- Fusion ---
merged = layers.Concatenate()([video_features, landmark_features])  # shape: (B, 608)
x = layers.Dense(256, activation='relu')(merged)
x = layers.Dropout(0.3)(x)
outputs = layers.Dense(num_classes, activation='softmax')(x)

model = tf.keras.Model(inputs=[video_input, landmark_input], outputs=outputs)
return model

I build this model but when I run model.fit() it shows in extract_video_features(x)
6 def extract_video_features(x):
7 endpoints = backbone(x, training=False)
----> 8 x = endpoints['head'] # shape: (B, T, 1, 1, C)
9 x = tf.squeeze(x, axis=[2, 3]) # shape: (B, T, C)
10 x = tf.keras.layers.GlobalAveragePooling1D()(x) # shape: (B, C)

TypeError: Exception encountered when calling Lambda.call().

tuple indices must be integers or slices, not str

Arguments received by Lambda.call():
• inputs=tf.Tensor(shape=(None, 50, 224, 224, 3), dtype=float32)
• mask=None
• training=True.

How can I fix this or any other ways to apply landmarks?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in build_classifier at extract_video_features and inspect what backbone(x, training=False) returns before endpoints['head'] is accessed; compare that return type and shape with the subsequent squeeze and pooling steps. Then validate the two-input model through model.fit using video and landmark inputs, with completion indicated by training running and producing outputs for num_classes.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.