NVIDIA-Merlin / NVIDIA-Merlin/Merlin

[RMP] Dynamic Batching support at serving time

Open
#906 0 comments 0 reactions 1 assignee View on GitHub

@EvenOldridge is already working on this.

Since Apr 12, 2023.

roadmap
Dominant language
Python
Stars
907
Forks
129
PR merge metrics
No merged PRs in 30d

Description

Problem:

Customers with high volumes of traffic want to trade off latency for throughput by grouping requests as dynamic batches.

Goal:

Leverage Triton's dynamic batching capabilities to enable support for dynamic batches in Merlin.

New Functionality

  • Models

    • ...
  • Transformers4Rec

    • ...
  • NVTabular

  • Dataloader

Systems
  • Dynamic batching with Triton
  • Serving-time padding operator (to use with dynamic batching)
Examples
  • Example of dynamic batching
  • Blog post on dynamic batching and tradeoff between latency and throughput.

Constraints:

Within Triton

Starting Point:

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.