prometheus / prometheus/client_python
Add an option to expose the metric timestamps from the Prometheus format
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 4.4k
- 派生
- 876
- 平均合并
- 8 天 4 小时
- 30 天内合并 PR
- 1
描述
Spin off from https://github.com/prometheus/client_python/pull/967 and https://github.com/prometheus/client_python/issues/847
Background
The Prometheus exposition format (https://github.com/prometheus/docs/blob/main/content/docs/instrumenting/exposition_formats.md#comments-help-text-and-type-information) allows exporters to specify an optional timestamp for each sample. If this is unset, the collector uses the timestamp that the sample is collected.
This timestamp is useful in a certain situation. My concrete use case is to expose GitHub API rate limit as a metric. This API rate limit is a value that decreases over time, and periodically it resets to the maximum value. We can get this API rate limit value as a HTTP response header when calling GitHub API.
Imagine a situation where there are two (physical) servers making this GitHub API call, and for the sake of illustration, let's assume that one server issues more requests and the other one does less.
| Server1 (frequently issue requests) | Server2 (somehow less requests) | |
|---|---|---|
| Observed API limit | 100 | 4900 |
| (Last time a server made a GH API call) | 3 minutes ago | 50 minutes ago |
From human's point of view, because server1 made an API call more recently than server2, we can tell that server1's observed API limit is the current value. However, without exposing the sample timestamp, Prometheus cannot distinguish which one is the most recent value. This is because when Prometheus collects sample from these two servers, when there's no timestamp specified, it treats these samples as "the samples collected now".
By exposing the sample timestamp, Prometheus should be able to treat these two samples correctly, and it can tell that server1' s value is the most recent value.
Request
With https://github.com/prometheus/client_python/pull/967, client_python should learn timestamps for Gauge metrics in the multiprocessing mode. This issue is asking for an option to expose those timestamps via Prometheus format.
Code pointers
Sample already can take a timestamp (https://github.com/prometheus/client_python/blob/249490e4ed30d4266182cfce50fe7046a484affd/prometheus_client/samples.py#L49), and so as the Prometheus exposition formatter (https://github.com/prometheus/client_python/blob/249490e4ed30d4266182cfce50fe7046a484affd/prometheus_client/exposition.py#L191-L194).
For the multiprocessing mode, changing the metrics aggregation logic to propagate the timestamps (https://github.com/prometheus/client_python/blob/249490e4ed30d4266182cfce50fe7046a484affd/prometheus_client/multiprocess.py#L146) would suffice (based on a flag or something).
For the non-multiprocessing mode, need to change this _child_samples (https://github.com/prometheus/client_python/blob/249490e4ed30d4266182cfce50fe7046a484affd/prometheus_client/metrics.py#L433-L434) to take a timestamp.
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 prometheus_client/samples.py 和 exposition.py 开始,追踪 sample 时间戳的表示和格式化方式。然后检查 metrics.py::_child_samples 和 multiprocess.py 中与所述聚合逻辑相关的部分,包括相关的 pull request 和 issue。完成的标准是:有一个选项可以在 Prometheus 输出中公开时间戳,并在 multiprocessing 和非 multiprocessing 模式下都保留这些时间戳。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- prometheus, python
- 领域
- observability
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 42/100