tensorflow / tensorflow/recommenders

[Question] How best to re-use large, repeated text features.

Open
#184 10 comments 0 reactions 1 assignee View on GitHub

@maciejkula is already working on this.

Since Dec 14, 2020.

question
Dominant language
Python
Stars
2k
Forks
300
PR merge metrics
No merged PRs in 30d

Description

Hi,

From reading the tutorials, the way data is passed into the model would mean lots of repetition, for example in the movie example, you would keep passing the same movie many times for different users. If we have a movie summary, this would mean having to store that summary many times over - which can get pretty big.

user | movie | movie summary

My question is what is the best way to create a "FeatureLookup" of sorts where for the same movie name / id we could grab the summary which would be stored once, and pass this in during training / inference?

tf.keras.Sequential([
  tf.keras.layers.experimental.preprocessing.StringLookup(
    vocabulary=unique_movies, mask_token=None
  ),
  tf.keras.layers.FeatureLookup(), # Fake layer, here as example, returns movie summary.
])

Thanks in advance!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.