prometheus / prometheus/client_python

Duplicated timeseries in CollectorRegistry with Multiprocess Gunicorn

オープン
#815 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

主要言語
Python
スター
4.4k
フォーク
876
平均マージ
8日 4時間
マージ済み PR(30日)
1

説明

I know this is a subject that comes up somewhat frequently but for the love of me I can't figure out what I'm doing wrong.

  1. I have a service in Amazon ECS thats running a single task with multiple workers (actually the problem happens in my other service that just has one worker also).

  2. I've created the directory and set the PROMETHEUS_MULTIPROC_DIR in the Dockerfile:

RUN mkdir -p /tmp/prom-metrics
ENV PROMETHEUS_MULTIPROC_DIR /tmp/prom-metrics
  1. I'm using the sample code in the README to create the registry in the /metrics request and return it:
registry = CollectorRegistry()
if getenv('PROMETHEUS_MULTIPROC_DIR'):
  multiprocess.MultiProcessCollector(registry)
data = generate_latest(registry)
status = '200 OK'
response_headers = [
    ('Content-type', CONTENT_TYPE_LATEST),
    ('Content-Length', str(len(data))),
]
return Response(data, status, response_headers)
  1. I've created the gunicorn.conf.py file with the sample from the README and passed it into my gunicorn startup script via -c:
from prometheus_client import multiprocess

def child_exit(server, worker):
    multiprocess.mark_process_dead(worker.pid)

In my two services, gunicorn starts them as follows:

# app 1 with workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 --workers 3 app:app

# app 2 without workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 app:app

The service boots successfully and accepts some metrics which are definitely collected in multiprocess mode, seeing as the HELP line simply displays Multiprocess metric.

This works for a few calls but eventually I get the dreaded Duplicated timeseries in CollectorRegistry error and no additional metrics are populated.

What might I be doing wrong?

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

README の multiprocess metrics の例と、参照されている gunicorn.conf.py 設定から始め、次に示されている Dockerfile の設定と Gunicorn コマンドを使って失敗を再現します。/metrics が引き続きメトリクスを返すように、重複した timeseries エラーの原因を特定し、設定を文書化または修正できれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
docker, prometheus, python
領域
observability
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。