[Feature Request]: Add metrics definitions to RunInference documentation
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
### What would you like to happen?
Add the following metrics definitions to RunInference documentation.
num_inferences
- The cumulative count of all samples being passed to RunInference. I.e. total sum of examples across all batches. This should increase monotonically.
inference_request_batch_size
- The number of samples in a particular batch of examples (created from beam.BatchElements) to be passed to run_inference(). This will vary over time depending on the dynamic batching decision of BatchElements().
inference_request_batch_byte_size
- The size, in bytes, of all elements for all samples in a particular batch of examples (created from beam.BatchElements) to be passed to run_inference(). This will vary over time depending on the dynamic batching decision of BatchElements(), and the particular values/dtypes of the elements.
inference_batch_latency_micro_secs
- The time, in microseconds, that it takes to perform the inference on the batch of examples. i.e. the time to call model_handler.run_inference(). This will vary over time depending on the dynamic batching decision of BatchElements(), and the particular values/dtypes of the elements.
model_byte_size
- The size, in bytes, of memory that the model takes for loading and initialization. i.e. the increase in memory usage from calling model_handler.load_model()
load_model_latency_milli_secs
- The time, in milliseconds, that it takes to load and initialize the model. i.e. the time it takes to call model_handler.load_model()
### Issue Priority
Priority: 2
### Issue Component
Component: run-inference
Contributor guide
Assessment
This issue has not been assessed yet.