prometheus / prometheus/client_python

[Question] Why isn't MultiProcessCollector a subclass of Collector?

未关闭
#1,105 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

主要语言
Python
星标
4.4k
派生
876
平均合并
8 天 4 小时
30 天内合并 PR
1

描述

I am studying how to use Prometheus in multiprocessing python code (without Gunicorn).

For now i came to following simple example with two child processes:


from multiprocessing import Process
import shutil
import time, os

os.environ["PROMETHEUS_MULTIPROC_DIR"] = './PMDir'  # This environment variable must be set BEFORE the first import 
                                                    # from prometheus_client to ensure the library uses MultiProcessValue 
                                                    # (metric value class with mmap support). Otherwise, 
                                                    # MutexValue (non-mmap) will be used, and files won't be created.

from prometheus_client import start_http_server, multiprocess, CollectorRegistry, Counter

os.environ["PROMETHEUS_MULTIPROC_DIR"] = './PMDir'  # If we set it AFTER import, the library initializes 
                                                    # the value class = MutexValue, mmap files will not be created,
                                                    # and MultiProcessCollector will have nothing to work with.
                                                    # This matches the documentation's recommendation to set it 
                                                    # before app startup.

COUNTER1 = None
COUNTER2 = None
COUNTER3 = None
COUNTER4 = None


def init_counters(registry):
    """
    Initialize counters (or other metrics) and register them in the specified registry.
    According to MetricWrapperBase's code, if registry=None, metrics won't be registered in any registry.
    """
    global COUNTER1, COUNTER2, COUNTER3, COUNTER4

    COUNTER1 = Counter('counter1', 'Incremented by the first child process', registry=registry)
    COUNTER2 = Counter('counter2', 'Incremented by the second child process', registry=registry)
    COUNTER3 = Counter('counter3', 'Incremented by all processes', registry=registry)
    COUNTER4 = Counter('counter4', 'Incremented by main process', registry=registry)

# We are free not to create registry object in child processes. Both f1 and f2 works as process targets.
# The mmap file handling is managed at the metric object level, not the collector level.
# Variation 1: create registry or not to create registry - both works as I expect.
def f1():
    """First child process body. Works without manual registry creation."""
    init_counters(None)
    while True:
        time.sleep(1)
        print("Child process 1", os.getpid())
        COUNTER1.inc()
        COUNTER3.inc()
    

def f2():
    """Second child process body. Works with manual registry creation."""
    registry = CollectorRegistry()
    init_counters(registry)
    while True:
        time.sleep(2)
        print("Child process 2", os.getpid())
        COUNTER2.inc()
        COUNTER3.inc()


if __name__ == '__main__':
    # Ensure the multiprocess directory exists and is empty
    prome_stats = os.environ["PROMETHEUS_MULTIPROC_DIR"]
    if os.path.exists(prome_stats):
        shutil.rmtree(prome_stats)
    os.mkdir(prome_stats)

    # Variation 2: When using MultiProcessCollector directly (see Variation 4), registry creation is optional
    # registry = CollectorRegistry()  # Create registry for HTTP server
    registry = None  # Works without registry between mpc and http server

    # Create MultiProcessCollector object. It reads and aggregates mmap files from PROMETHEUS_MULTIPROC_DIR.
    # Registering it in our registry means thar registry.collect() calls mpc.collect() and thus metrics from 
    # mpc aggregated by registry.
    # MultiProcessCollector ONLY reads (mmap) files; metric saving is handled by MultiProcessValue in metrics obj.
    mpc = multiprocess.MultiProcessCollector(registry)

    # Variation 3: If main process have to report its own metrics it can use both separate CollectorRegistry or None.
    # init_counters(CollectorRegistry())  # Use separate registry for main process metrics.
    init_counters(None)  # Works without registry registration
    # init_counters(registry)  # Metrics will duplicate if we use same registry as that one mpc registered in.
    #                          # More precisely, in this situation registry will export both mpc metrics and 
    #                          # main process metrics, including overlaps

    # Variation 4: HTTP server can use both registry or mpc
    # start_http_server(8000, registry=registry)  # Standard approach
    start_http_server(8000, registry=mpc)  # Works despite mpc not being a Collector subclass (but implements collect())

    p1 = Process(target=f1, args=())
    p1.start()
    p2 = Process(target=f2, args=())
    p2.start()

    print("collect")

    try:
        while True:
            print('main process   ', os.getpid())
            time.sleep(1)
            COUNTER3.inc()
            COUNTER4.inc()
    except KeyboardInterrupt:
        p1.terminate()
        p2.terminate()
        shutil.rmtree(prome_stats)

(I left my findings as comments because documentation didnt answer my questions and I can make mistakes in my assumptions)

I noticed that MultiProcessCollector implements collect() but doesn't inherit from the Collector class.

The Custom collectors documentation suggests that custom collectors should implement this interface. However, MultiProcessCollector is defined as:

class MultiProcessCollector:
    """Collector for files for multi-process mode."""

    def __init__(self, registry, path=None):
        ...

This seems contradictory since:

  • The class name contains "Collector"
  • It implements the core collect() method
  • It's designed to be registered in a CollectorRegistry

Could you clarify:

  1. Is this intentional design? If yes, what's the rationale behind not inheriting from Collector?
  2. Are there any potential compatibility risks in not following the Collector interface contract?
  3. What are potential risks of using MultiProcessCollector object in start_http_server directly, without registry = CollectorRegistry() between them?

I want to ensure I understand this correctly for proper integration. Thanks for your insights!

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

先从链接的 Custom collectors 文档开始,然后检查 client_python 仓库中所示的 MultiProcessCollector 和 start_http_server 入口点。将它们的注册和收集行为与文档化的 Collector 接口进行比较。当接口设计依据、兼容性注意事项以及直接使用 start_http_server 的方式都得到清晰记录时,即可视为完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
observability
Issue 类型
文档
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。