prometheus / prometheus/client_python
Duplicated timeseries in CollectorRegistry with Multiprocess Gunicorn
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 4.4k
- 派生
- 876
- 平均合并
- 8 天 4 小时
- 30 天内合并 PR
- 1
描述
I know this is a subject that comes up somewhat frequently but for the love of me I can't figure out what I'm doing wrong.
-
I have a service in Amazon ECS thats running a single task with multiple workers (actually the problem happens in my other service that just has one worker also).
-
I've created the directory and set the
PROMETHEUS_MULTIPROC_DIRin the Dockerfile:
RUN mkdir -p /tmp/prom-metrics
ENV PROMETHEUS_MULTIPROC_DIR /tmp/prom-metrics
- I'm using the sample code in the README to create the registry in the
/metricsrequest and return it:
registry = CollectorRegistry()
if getenv('PROMETHEUS_MULTIPROC_DIR'):
multiprocess.MultiProcessCollector(registry)
data = generate_latest(registry)
status = '200 OK'
response_headers = [
('Content-type', CONTENT_TYPE_LATEST),
('Content-Length', str(len(data))),
]
return Response(data, status, response_headers)
- I've created the
gunicorn.conf.pyfile with the sample from the README and passed it into my gunicorn startup script via-c:
from prometheus_client import multiprocess
def child_exit(server, worker):
multiprocess.mark_process_dead(worker.pid)
In my two services, gunicorn starts them as follows:
# app 1 with workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 --workers 3 app:app
# app 2 without workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 app:app
The service boots successfully and accepts some metrics which are definitely collected in multiprocess mode, seeing as the HELP line simply displays Multiprocess metric.
This works for a few calls but eventually I get the dreaded Duplicated timeseries in CollectorRegistry error and no additional metrics are populated.
What might I be doing wrong?
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
先从 README 中的 multiprocess metrics 示例和所引用的 gunicorn.conf.py 配置开始,然后使用所示的 Dockerfile 设置和 Gunicorn 命令重现该故障。完成的标准是确定 duplicated timeseries 错误的原因,并记录或修正配置,使 /metrics 继续返回 metrics。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- docker, prometheus, python
- 领域
- observability
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 需要澄清
- 新手友好度
- 25/100