grpc / grpc/grpc-java

netty: Writes can starve reads, causing DEADLINE_EXCEEDED

Open
#8,912 6 comments 0 reactions 1 assignee Claimed by @sergiitk View on GitHub
netty performance
Dominant language
Java
Stars
12.1k
Forks
4k
Avg merge
2d 17h
Merged PRs (30d)
37

Description

### What version of gRPC-Java are you using?
`1.44.0`

### What is your environment?

A client and server communicating over localhost.

Operating system: `Mac OSX`
JDK: `11.0.9`

### What did you expect to see?
Client in this [branch](https://github.com/tommyulfsparre/grpc-half-close/tree/netty-okhttp-slow-start) not failing the majority of RPCs with `DEADLINE_EXCEEDED` when using Netty for transport.

### What did you see instead?

The majority of call fails with `DEADLINE_EXCEEDED` using `Netty` whereas using `OkHttp` doesn't .

### Steps to reproduce the bug

This (hacky) [reproducible](https://github.com/tommyulfsparre/grpc-half-close/tree/netty-okhttp-slow-start) might indicate a potential problem where the periodic flushing of the [WriteQueue](https://github.com/grpc/grpc-java/blob/master/netty/src/main/java/io/grpc/netty/WriteQueue.java#L122) doesn't happen frequently enough. This causes latency leading to `DEADLINE_EXCEEDED` during startup. Switching to `ÒkHttp` for **this** particular case does not cause any RPCs to fail.

On the wire with Netty multiple RPCs are batched and sent togheter whereas with `ÒkHttp` they are immediately dispatched.

Here is a screenshot from Perfmark using Netty:

![permark](https://user-images.githubusercontent.com/2462159/153612117-ab742c54-fc11-4702-9ce5-e18b91d1891c.jpg)

is this a known issue or am I doing something fundamentally wrong in the reproducer?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.