tensorflow / tensorflow/recommenders

Some issues not covered in the tutorials

Open
#190 8 comments 0 reactions 1 assignee View on GitHub

@maciejkula is already working on this.

Since Dec 21, 2020.

question
Dominant language
Python
Stars
2k
Forks
300
PR merge metrics
No merged PRs in 30d

Description

Apologies, a noob here but I suspect my questions will be repeated by other noobs :)

Could the developers please elaborate on the below (I think these could be good additions to the tutorials too).

  1. How can one specify the number of recommendations to generate per user? The default appears to be 10; where can this be overridden?
  2. How does the caller retrieve the generated recommendations along with the respective recommendation ratings?
    _, titles = index(tf.constant(["42"]))
    print(f"Recommendations for user 42: {titles[0, :3]}")
    This retrieves the movie titles, in the tutorial, but not the rating values.
    (OK, x, titles = index(tf.constant(["42"])) -- looks like the x tensor contains the ratings)
  3. The generation of the embedding values e.g. for user ID's. Must these be contiguous integers? Can I re-use my own ID values? E.g. I have two users, user 1 with ID=123 and user 2 with ID=998. Can I use 123 and 998 or must I map these ID's to 1 and 2? (I assume it's the latter approach but please clarify).
  4. Is there a way to instruct the recommenders code not to include user's history items in the generated recommenders? The idea is to avoid having to do my own post-filter.
  5. Is there a way to instruct the recommenders code not to include multiples of the same item in the generated recommendations? There was a sentence in the tutorial that led me to believe that duplicates might be present (please clarify).
  6. During featurization/tokenization, have folks worked out the I18N aspects? Say, if I have text features in English and Spanish, are both an English tokenizer and a Spanish tokenizer available? How well would they work with short strings? Is there a Language Identifier to wire in?
  7. What is a 'good' range of values for the top-100 accuracy? should it be close to 1? how close?
  8. How to bring RMSE down? When running some of the examples, RMSE tends to be > 1. Is there an optimal number of features to use, perhaps? Any suggestions as to the tuning of the hyperparameters to keep the RMSE below 1?
  9. How to scale/distribute the processing? If I have several million users and several million items and 3-4 features on users and 5-7 features on the items, what would be some of the approaches to scale this? I'm looking at this writeup: https://towardsdatascience.com/scaling-up-with-distributed-tensorflow-on-spark-afc3655d8f95. Any examples of how to code up a TFRS recommender that could run on Spark? or some other way to distribute. Seems like TensorFlowOnSpark is a way to go...
  10. How to tone down the amount of prints in the console.

Thanks.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.