prometheus / prometheus/prometheus

Use metric_relabel_configs to drop metrics but still get a high prometheus memory usage

Open
#13,836 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Go
Stars
66.1k
Forks
10.8k
Avg merge
2d 1h
Merged PRs (30d)
131

Description

What did you do?

We have the prometheus setup and it scapes metrics from service A, service A has 450k time series, and the prometheus memory usage is 2.6GB.

Now we want promethues scape metrics from both service A and service B, service B has 1300k time series, and we set below metric_relabel_configs setting for service B, it means that we will only keep the time series which contains myMetric at its label. Now we don't have any metric contains myMetric label.

If we check the metric scrape_samples_scraped, the number of time series is 450K+1300K
If we check the metric scrape_samples_post_metric_relabeling, the number of time series is 450k

  # scrape from service B
  - job_name: "serviceB"
    scrape_interval: 20s
    metrics_path: "xxx"
    static_configs:
      - targets: ["localhost:xxx"]
    metric_relabel_configs:
      - source_labels: ["myMetric"]
        regex: ".+"
        action: keep
What did you expect to see?

We expect to see the same prometheus memory usage 2.6GB because metric_relabel_configs happened before the data is ingested by the storage system.

What did you see instead? Under which circumstances?

We see a prometheus memory increase from 2.6GB to 3.3GB.

System information

Linux 5.15.0-1059-azure x86_64

Prometheus version
v2.45.0
Prometheus configuration file

No response

Alertmanager version

No response

Alertmanager configuration file

No response

Logs

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the service A and service B scrape setup with the shown metric_relabel_configs rule, then compare scrape_samples_scraped with scrape_samples_post_metric_relabeling and Prometheus memory usage. Determine which stage accounts for the retained memory and document the behavior or required change; the issue provides no source file or test to target.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
observability-sre, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.