NVIDIA-Merlin / NVIDIA-Merlin/Transformers4Rec
[BUG] XLNET-CLM eval recall metric value does not match with custom np based recall metric value
Open
@marcromeyn is already working on this.
Since Jun 12, 2023.
bug
P1
- Dominant language
- Python
- Stars
- 1.3k
- Forks
- 165
- Avg merge
- 1m
- Merged PRs (30d)
- 2
Description
Bug description
When we train an XLNet model with CLM masking, the model prints out its own evaluation metrics (ndcg@k, recall@k, etc.) from trainer.evaluate() step. If we want to apply our own custom metric func using numpy something like below, the metric values do not match, but they match if we use MLM masking instead.
def recall(predicted_items: np.ndarray, real_items: np.ndarray) -> float:
bs, top_k = predicted_items.shape
valid_rows = real_items != 0
# reshape predictions and labels to compare
# the top-10 predicted item-ids with the label id.
real_items = real_items.reshape(bs, 1, -1)
predicted_items = predicted_items.reshape(bs, 1, top_k)
num_relevant = real_items.shape[-1]
predicted_correct_sum = (predicted_items == real_items).sum(-1)
predicted_correct_sum = predicted_correct_sum[valid_rows]
recall_per_row = predicted_correct_sum / num_relevant
return np.mean(recall_per_row)
Steps/Code to reproduce bug
coming soon.
Expected behavior
Environment details
- Transformers4Rec version:
- Platform:
- Python version:
- Huggingface Transformers version:
- PyTorch version (GPU?):
- Tensorflow version (GPU?):
Additional context
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.