NVIDIA-Merlin / NVIDIA-Merlin/Merlin

[RMP] Quick-start for retrieval models training pipeline

Open
#828 0 comments 0 reactions 2 assignees View on GitHub

@sararb is already working on this.

Since Feb 22, 2023.

documentation roadmap
Dominant language
Python
Stars
907
Forks
129
PR merge metrics
No merged PRs in 30d

Description

Problem:

Merlin provides documentation and a number of example notebooks on how to use tools like NVTabular, Dataloader and Merlin Models. In order to build a pipeline for training and evaluation purposes, a Data Scientist needs to analyze that material, copy-and-paste code snippets demonstrating the API and glue that code together to implement scripts for experimentation and benchmarking.
It might also not be clear to the users the advanced API options featured by Merlin Models that can be mapped as a hyperparameter, and potentially improve models accuracy.

Goal:

This RMP provides a Quick-start for building retrieval models training pipelines.
It addresses the retrieval models part of this larger RMP #732, in particular the steps 4-7 of the Data Scientist journey when experimenting with Merlin Models.

The Quick start for retrieval is composed by:

Template scripts
  • Generic template script for preprocessing
  • Generic template script for training retrieval models (MF, TwoTower, YouTubeDNN), exposing the main hyperparameters.
Documentation
  • Documentation of the scripts command line arguments
  • Documentation of best practices learned from our experimentation:
    • Hyperparameter tuning: search space, most important hyperparameters and best hparams for TenRec public dataset for STL and MTL models
    • Intuitions of API options (building blocks, arguments) that can improve models accuracy

Constraints:

  • Preprocessing - The pre-processing template notebook will perform some basic feature encoding for categorical (e.g. categorify) and continuous variables (e.g. standardization). The customer can expand the template with advanced preprocessing ops demonstrated in our examples.
  • Training - The training and evaluation script for Merlin Models should be totally configurable, taking as input the parquet files and schema, and a number of hyperparameters exposed via command line arguments. The output of this script should be the evaluation metrics, logged as a CSV file and also to Weights&Biases.

Starting Point:

The retrieval training scripts we have developed for the Retrieval and Two-stage research projects.

Tasks:

Tenrec dataset experiments
  • NVIDIA-Merlin/models#803
Research & Documentation
  • NVIDIA-Merlin/models#802
  • NVIDIA-Merlin/models#804
Testing
  • NVIDIA-Merlin/models#805

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.