libp2p / libp2p/test-plans

Benchmarking feedback/notes

Open
#222 4 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
65
Forks
82
PR merge metrics
No merged PRs in 30d

Description

Spent a bit of time looking at some parts of the benchmarking setup, and had a couple of notes and comments:

* I think we're using iperf wrong. We are using the data from the sender, we should be looking at the receiver. Notice how the Bitrate is very different for sender vs receiver in this example:
```
$ iperf -c 127.0.0.1 -u -b 10g
Connecting to host 127.0.0.1, port 5201
[ 5] local 127.0.0.1 port 50191 connected to 127.0.0.1 port 5201
[ ID] Interval Transfer Bitrate Total Datagrams
[ 5] 0.00-1.00 sec 1.16 GBytes 10.0 Gbits/sec 38132
[ 5] 1.00-2.00 sec 1.16 GBytes 10.0 Gbits/sec 38161
[ 5] 2.00-3.00 sec 1.16 GBytes 9.99 Gbits/sec 38106
[ 5] 3.00-4.00 sec 1.17 GBytes 10.0 Gbits/sec 38184
[ 5] 4.00-5.00 sec 1.16 GBytes 10.0 Gbits/sec 38151
[ 5] 5.00-6.00 sec 1.16 GBytes 10.0 Gbits/sec 38143
[ 5] 6.00-7.00 sec 1.16 GBytes 9.99 Gbits/sec 38114
[ 5] 7.00-8.00 sec 1.16 GBytes 10.0 Gbits/sec 38165
[ 5] 8.00-9.00 sec 1.16 GBytes 9.99 Gbits/sec 38140
[ 5] 9.00-10.00 sec 1.16 GBytes 10.0 Gbits/sec 38169
- - - - - - - - - - - - - - - - - - - - - - - - -
[ ID] Interval Transfer Bitrate Jitter Lost/Total Datagrams
[ 5] 0.00-10.00 sec 11.6 GBytes 10.0 Gbits/sec 0.000 ms 0/381465 (0%) sender
[ 5] 0.00-10.00 sec 8.43 GBytes 7.24 Gbits/sec 0.023 ms 105024/381411 (28%) receiver
```

We need to use the bitrate on the receiver side. The sender can push as much data as you want, but for these measurements we care about the data that was actually received. Look at the difference here: https://github.com/libp2p/test-plans/actions/runs/5466146370/jobs/9950640038#step:12:29
* The hypothetical max for this use case should be 50% of the instance bandwidth according to https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-instance-network-bandwidth.html. (3.4 Gbps). I think it's worth linking this doc somewhere.
* The "local" vs "remote" backends are a bit confusing. These are both running on AWS hardware. Could we consolidate them or rename them? I would suggest alternate names, but I don't really understand them.
* What's the ami of the short-lived module? Doesn't seem set, and I can't find the default
* Should we make sure to set the MTU to 1500? (This might not be the default)
* Do we need to bump the UDP send window as well? I'm not sure, but it might be fine since quic-go doesn't complain about it. Any insight here @marten-seemann?
* Can we add comments around the AMI ids to describe them? It wasn't clear that these were the Amazon Linux AMIs
* Maybe include this one-liner:
```
aws ec2 describe-images \
--image-id ami-06e46074ae430fba6 \
--query "Images[*].Description[]" \
--output text \
--region us-east-1
```

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the benchmarking setup and the linked GitHub Actions run, then compare the measurements with the AWS instance network-bandwidth documentation. Investigate the local and remote backends, the short-lived module's AMI, MTU and UDP settings, and the comments around the listed AMI IDs. Done means the requested benchmark configuration and documentation questions have clear, agreed resolutions.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws
Domain
cloud, networking, performance
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.