ros2 / ros2/performance_test_fixture

Benchmarks with small ratio of code-under-test runtime to setup runtime can lead to long benchmark evaluation times

Open
#12 0 comments 0 reactions 1 assignee View on GitHub

@cottsay is already working on this.

Since Dec 3, 2020.

Dominant language
C++
Stars
11
Forks
7
Avg merge
20h 36m
Merged PRs (30d)
1

Description

Issue

Based on discussion over at https://github.com/ros2/rclcpp/pull/1447#discussion_r523162220

If setup code is needed during each iteration of a benchmark run, and that setup code can take a long time relative to the code under test, google benchmarks will choose to run the benchmark for an abnormally long time. This can happen when PauseTiming/ResumeTiming or SetIterationTime is used to avoid google benchmark from measuring the full iteration time.

For example, this benchmark was removed in https://github.com/ros2/rclcpp/pull/1447 due to it's long runtime. (Note this version is fixed so the destructors of client/service are not included in the measured time).

BENCHMARK_F(PerformanceTest, get_node_name)(benchmark::State & state)
{
 int count = 0;
  for (auto _ : state) {
    state.PauseTiming();
    const std::string service_name = std::string("service_") + std::to_string(count++);
    auto client = node->create_client<test_msgs::srv::Empty>(service_name);
    auto callback =
      [](
      const test_msgs::srv::Empty::Request::SharedPtr,
      test_msgs::srv::Empty::Response::SharedPtr) {};
    auto service =
      node->create_service<test_msgs::srv::Empty>(service_name, callback);
    state.ResumeTiming();

    client->wait_for_service(std::chrono::seconds(1));
  
    state.PauseTiming();
    client.reset();
    service.reset();
    state.ResumeTiming();
  }
}

Here client->wait_for_service() was measured at 4us, but creating the client, service and then destroying both would take many orders of magnitude more time. Google benchmarks chooses to run based on the configured MinTime, which is defaulted to about a second. So it's looking to run this benchmark about 200,000 iterations.

Based on these benchmark results:
http://build.ros2.org/view/Rci/job/Rci__benchmark_ubuntu_focal_amd64/193/testReport/projectroot.test/benchmark/benchmark_client/
http://build.ros2.org/view/Rci/job/Rci__benchmark_ubuntu_focal_amd64/193/testReport/projectroot.test/benchmark/benchmark_service/

Constructing and destructing a client and service both would take about 500us-600us. At 200,000 iterations, this benchmark would be required to run for 120 seconds to achieve 1 second of evaluation time for the code under test.

Possible resolutions

There may be more resolutions identified, these are just suggestions to get the ball rolling.

  1. Make use of MinTime and accept the reduced accuracy of measurement averages. Calling MinTime on a registered benchmark, can override the default value of 1 second. The MinTime duration would need to be chosen specifically for each benchmark and might need to be periodically updated if significant changes to the source code occur. The downside of reducing the MinTime for a benchmark is that fewer iterations are measured and that could decrease the accuracy of the averaged result.

  2. Measure the full iteration time with and without code under test. In the example above, it would require creating two benchmarks. One with the call to wait_for_service and another without it. The benchmarks could then be compared side-by-side. The benefit here is that the reduced number of iterations is automatically chosen by google benchmark and can adjust as the source code changes. The disadvantage is that we'd be measuring the difference of averages, which would be more noisy than choice 1.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.