mlcommons / mlcommons/modelbench

Why was nvidia-llama-3-1-nemotron-nano-8b-v1 run so slow?

Open
#989 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
134
Forks
36
Avg merge
1d 11h
Merged PRs (30d)
17

Description

This run took more than 6 hours, but it only needed to run two items:

sut_cache: DiskCache(run/sut_cache)
sut_cache: started with 11997
sut_cache: finished with 11997
annotator_cache: DiskCache(run/annotator_cache)
annotator_cache: started with 47340
annotator_cache: finished with 47318

And according to the cache info, it did slightly less than nothing:

sut_cache: DiskCache(run/sut_cache)
sut_cache: started with 11997
sut_cache: finished with 11997
annotator_cache: DiskCache(run/annotator_cache)
annotator_cache: started with 47340
annotator_cache: finished with 47318

More signs that something about the cache isn't right.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the nvidia-llama-3-1-nemotron-nano-8b-v1 run and inspect the run/sut_cache and run/annotator_cache entries and counts. Trace why the run exceeds six hours and why annotator_cache decreases from 47340 to 47318; done means the cache behavior and runtime cause are explained and the issue has a verified fix.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.