apache / apache/pulsar

PIP-202: Support latency quantile metric for pulsar read and write

Open
#17,090 5 comments 0 reactions 0 assignees View on GitHub
Stale type/PIP
Dominant language
Java
Stars
15.3k
Forks
3.8k
Avg merge
1d 14h
Merged PRs (30d)
160

Description

### Motivation

1. Currently, pulsar does not have an quantile metric used to statistics read and write latency, which is essential for our future troubleshooting and performance comparison tests.
2. And we need network io metrics to determine the current network processing pressure, to prevent too many requests from hanging the broker.

### Goal

1. Our goal is to statics the latency which begin at readHandler init and end at the messages are read from the cache or underlying storage(eg: bookkeeper / hdfs / ...).
Why we just statics the latency between readHandler init and messages are found:
1.1. Because when fetch request arrive the broker, It may be no new messages for Consumer, the broker will wait and periodically query whether there are new messages, so it may cause the read cache latency longer than the remote read.
1.2. We statics the read latency between the cache/bookeeper/hdfs..., it is enough for us to troubleshooting and performance comparison tests.

2. We just statics the idle network process number and its percnetile of the total number of io threads.

### API Changes

1. Add OpStats get method to statics the write latency:
```
public OpStatsLogger getOpStat();
```
2. Add generate metrics method for netty thread pool usage
```
private static void generateNetworkIdleMetrics(PulsarService pulsar, SimpleTextOutputStream stream);
```

### Implementation

1. Add OpStats get method to statics the write latency:
```
public OpStatsLogger getOpStat() {
return PulsarService.statsProvider.getStatsLogger("").getOpStatsLogger("READ_ENTRY_LATENCY");
}
```

2. Add generate metrics method for netty thread pool usage:
```
private static void generateNetworkIdleMetrics(PulsarService pulsar, SimpleTextOutputStream stream) {
// generate network idle percent metrics
try {
int busyExecutors = 0;
EventLoopGroup workerGroup = pulsar.getBrokerService().executor();
Iterator iterator = workerGroup.iterator();
while (iterator.hasNext()) {
SingleThreadEventExecutor next = (SingleThreadEventExecutor) iterator.next();
if (next.pendingTasks() > 0) {
++busyExecutors;
}
}
int numIoThreads = pulsar.getConfiguration().getNumIOThreads();
float ioWaitRatioMetric = (float) busyExecutors / (float) numIoThreads;
// Metric: netIdlePercentile -> ioWaitRatioMetric
writeNetIdlePercentileMetrics(stream, "brk_net_idle_percentile", (1 - ioWaitRatioMetric) * 100f,
Collector.Type.GAUGE, clusterName, currentTimeMillis);

// general network io queue metrics
Iterator iterator1 = workerGroup.iterator();
while (iterator1.hasNext()) {
SingleThreadEventExecutor next = (SingleThreadEventExecutor) iterator1.next();
int pendingTasks = next.pendingTasks();
String name = "brk_pending_tasks_" + StringUtils.replace(next.threadProperties().name(),
"-", "_");
// Metric: thread-name -> pendingTasks
writeNetIdlePercentileMetrics(stream, name, pendingTasks,
Collector.Type.GAUGE, clusterName, currentTimeMillis);
}
} catch (Exception e) {
log.error("generate network idle percent metrics failed, error: [()]", e);
}
}
```

Contributor guide

Open the contributing guide

Research direction

Locate the existing OpStatsLogger API and the PulsarService metrics-generation entry point first, then trace how broker latency and network metrics are currently exposed. Use the proposed getOpStat and generateNetworkIdleMetrics snippets to define the scope; done means the requested read/write latency and idle or pending-task metrics are exposed through broker metrics output.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
backend, observability
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.