grpc / grpc/grpc-rust

Question: lighter unary path for tiny high-rate RPCs

Open
#2,832 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
12.5k
Forks
1.3k
Avg merge
4d 7h
Merged PRs (30d)
24

Description

## Question

For tiny high-rate unary RPCs, is there an actionable optimization point in tonic's unary pipeline, or an intended "lighter unary" path that keeps gRPC compatibility while avoiding some per-call fixed cost?

This is not a bug report against correctness. I am trying to decide whether an application should keep optimizing around batching/coalescing RPCs, or whether there is useful upstream work in tonic/h2 for small unary calls.

## Minimal repro

Repro gist: https://gist.github.com/houseme/ab26954bf2600a907c14e873a3332fd0

Run:

```console
cargo run --release
```

The repro compares four paths on loopback with a 1 KiB response payload and 20,000 calls:

- `raw_h2_prost`: direct h2 request/response with prost payload, no gRPC frame/status semantics.
- `raw_h2_grpc_like`: direct h2 with `content-type: application/grpc`, `te: trailers`, 5-byte gRPC message prefix, response data, and `grpc-status: 0` trailers.
- `tonic_unary_plain`: `tonic::client::Grpc::unary` against a minimal tonic server service and prost codec.
- `tonic_unary_minimal_tower_wrapper`: same as tonic plain, with one minimal tower `map_request` wrapper to approximate application middleware shape.

## Results

Environment:

- rustc 1.98.0 (88d9e12ae 2026-08-18)
- cargo 1.98.0 (797e8a9bc 2026-08-05)
- host: aarch64-apple-darwin
- tonic 0.14.6
- h2 0.4.18
- prost 0.14.4
- bytes 1.12.1

Run 1:

```text
raw_h2_prost c=1: 54.05 us/call
raw_h2_prost c=64: 8.37 us/call
raw_h2_grpc_like c=1: 77.37 us/call
raw_h2_grpc_like c=64: 10.66 us/call
tonic_unary_plain c=1: 97.30 us/call
tonic_unary_plain c=64: 17.92 us/call
tonic_unary_minimal_tower_wrapper c=1: 99.59 us/call
tonic_unary_minimal_tower_wrapper c=64: 18.30 us/call
```

Run 2:

```text
raw_h2_prost c=1: 53.58 us/call
raw_h2_prost c=64: 8.20 us/call
raw_h2_grpc_like c=1: 77.99 us/call
raw_h2_grpc_like c=64: 10.48 us/call
tonic_unary_plain c=1: 97.53 us/call
tonic_unary_plain c=64: 18.00 us/call
tonic_unary_minimal_tower_wrapper c=1: 95.69 us/call
tonic_unary_minimal_tower_wrapper c=64: 17.93 us/call
```

The intentionally incomplete raw h2/prost path is only a lower bound. The more relevant comparison is `raw_h2_grpc_like` versus `tonic_unary_plain`, where this repro still shows roughly:

- Sequential: about +19 to +20 us/call for tonic unary.
- Concurrency 64: about +7 to +8 us/call for tonic unary.
- The minimal extra tower wrapper is near noise in this repro, so the remaining gap appears more likely around tonic unary's client/server pipeline, codec/body/status/trailer handling, or readiness/service machinery than around one application wrapper.

## Context

This came from a RustFS investigation of many small per-object internal metadata RPCs. In the application profile, h2/writev and tonic client-side framing/header work dominate relative to the tiny payload. RustFS-specific auth/header construction and wrapper allocation were checked separately and did not currently look like the best optimization target.

We also tested HTTP/2 stream/connection window changes separately in the application and did not see a meaningful improvement for these 1 KiB metadata RPCs, which makes sense because this looks like fixed per-RPC overhead rather than BDP/window pressure.

## What I am asking

- Is this benchmark shaped reasonably enough to discuss tonic unary fixed overhead, or is it missing an important generated-client/server behavior that would materially change the result?
- Is there an existing supported lower-overhead unary path for tiny internal RPCs that still keeps normal gRPC compatibility?
- If not, are there known hotspots in tonic's unary pipeline worth profiling further or turning into a PR, for example around per-call header/trailer/status handling, body construction, codec dispatch, readiness, or boxing?

If the answer is that this is mostly unavoidable gRPC/tonic semantic cost, that is also useful: in that case we will shift the RustFS optimization direction toward reducing per-object RPC count via request coalescing/batching rather than trying to shave small per-RPC allocations locally.

Contributor guide

Open the contributing guide

Research direction

Start with the tonic::client::Grpc::unary path and the minimal tonic server service described in the repro, then compare them with the raw h2_grpc_like path. Profile the client/server pipeline, codec and body handling, headers, trailers, status processing, readiness, and boxing; done means documenting whether a compatible lower-overhead path or a focused upstream optimization is justified.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend-api-design, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.