NVIDIA-Merlin / NVIDIA-Merlin/Merlin
[RMP] Dynamic Batching support at serving time
Open
@EvenOldridge is already working on this.
Since Apr 12, 2023.
roadmap
- Dominant language
- Python
- Stars
- 907
- Forks
- 129
- PR merge metrics
- No merged PRs in 30d
Description
Problem:
Customers with high volumes of traffic want to trade off latency for throughput by grouping requests as dynamic batches.
Goal:
Leverage Triton's dynamic batching capabilities to enable support for dynamic batches in Merlin.
New Functionality
-
Models
- ...
-
Transformers4Rec
- ...
-
NVTabular
-
Dataloader
Systems
- Dynamic batching with Triton
- Serving-time padding operator (to use with dynamic batching)
Examples
- Example of dynamic batching
- Blog post on dynamic batching and tradeoff between latency and throughput.
Constraints:
Within Triton
Starting Point:
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.