tensorflow / tensorflow/recommenders

[Question] Difference between top_k_categorical_accuracy_at_100 and recall at 100

Open
#617 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2k
Forks
300
PR merge metrics
No merged PRs in 30d

Description

Hi Team,

Thanks for the awesome discussions, I've been learning a lot from the discussions in the issues, and many of them improved the accuracy of the work I'm doing.
I'm wondering if there's a big difference in calculating these two metrics or not as I'm able to get around 0.9 top_k_categorical_accuracy_at_100, but when I'm doing offline evaluation using recall at k I get around 0.35.

From my understanding top_k_categorical_accuracy_at_100 is count(interactions that the ground truth was ranked in the top k)

If my batch size is 1028, I can assume that top_k_categorical_accuracy_at_100 of 0.9 means that out of random 1028 candidates we get the right one in the top 100 candidates 90% of the cases, am i missing sth?

This is very close to the definition of recall I'm calculating
recall at 100 = intersection(user future interactions, top 100 candidates for this user embedding)/len(user future interactions)

The key differences are
for top_k_categorical_accuracy_at_100:

  1. candidates are selected randomly and I've a total of 3000 candidates, so many duplicates will appear in a batch of 1028 so the results would be optimistic (as the unique candidates would be much less than that)
  2. not all candidates (restaurants) can be deliver to queries (users) for all cases, so the right candidates are a subset of all candidates, if the model learnt that, it'll be able to easily neglect the candidates that can't interact with the query --> which will lead to an easier job for classifier.
    On the other hand

for Recall at 100 calculation

  1. I only considers the candidates (restaurants) that can deliver to the user so it's a harder problem
  2. No duplications, as i only deal with the unique candidates.

But I don't expect that the results would be that different 95% for top_k_categorical_accuracy_at_100 and 35% for recall at 100

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified in the issue. Start by examining the implementations and evaluation definitions for top_k_categorical_accuracy_at_100 and recall at 100; done means clearly explaining why the reported values differ and whether the comparison is valid.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.