apache / apache/paimon

[Bug] TaskManager Metaspace never reclaimed: ~18.2 MB leaked per finished SQL job that ships the Paimon connector as a user jar (Flink 1.20.5 session cluster)

Open
#9,795 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
3.4k
Forks
1.4k
Avg merge
1d 11h
Merged PRs (30d)
396

Description

### Search before asking

- [X] I searched in the [issues](https://github.com/apache/paimon/issues) and found nothing similar.

Closest existing reports, both still open, neither matching this measurement set: #2780 (`Flink S3 cause TaskManager Metaspace OOM`, Paimon 0.6.0 / Flink 1.18, attributes it to an unclosed `S3AFileSystem`) and #5998 (`paimon-s3` vs `flink-s3-fs-hadoop` plugin conflict). On the Flink side: FLINK-28248 (open, no fix version, but the reproducer and root cause there are Beam-specific), FLINK-19005 (closed by documenting JDBC drivers as a cause — excluded here, see below), FLINK-32203 (log4j `monitorInterval` — excluded here, see below).

### Paimon version

Paimon Flink connector 1.2.0 / 1.3.1 (`paimon-flink-1.20`, 52.8 MB shaded jar), `paimon-s3-1.3.1.jar`.

### Compute Engine

Apache Flink **1.20.5**, standalone **session cluster** on Kubernetes.
JVM: Temurin **17.0.19** (the official Flink image ships a **JRE**), G1GC.
TaskManager `-XX:MaxMetaspaceSize=768m`, JobManager `2048m`.
Storage: S3 (MinIO), catalog declared as `s3://warehouse/paimon` with `s3.endpoint` / `s3.path.style.access` (Paimon-native S3, deliberately **not** `s3a://` + `fs.s3a.*`).

### Minimal reproduce step

In a long-running session cluster, submit this repeatedly (each invocation is a separate short batch job; the Paimon jar is shipped as a **user jar** via `-j`):

```sql
-- q.sql
SET 'pipeline.name' = 'probe';
SET 'execution.runtime-mode' = 'batch';
CREATE CATALOG lake WITH (
'type'='paimon',
'warehouse'='s3://warehouse/paimon',
's3.endpoint'='http://minio:9000',
's3.path.style.access'='true',
's3.access-key'='...','s3.secret-key'='...');
SELECT COUNT(*) FROM lake.`mydb`.`small_pk_table`;
```

```bash
/opt/flink/bin/sql-client.sh -j /path/paimon-sql.jar -f q.sql
```

Table under test: Paimon **primary-key** table, `parquet` + `zstd`, `bucket=4`, 24 columns, 11 rows. All jobs finish normally (`FINISHED`, correct result).

### What doesn't meet your expectations?

**Each finished job permanently costs ~18.2 MB of TaskManager Metaspace, and it is never returned.**

Controlled measurement (cluster-wide sum of `Status.JVM.Memory.Metaspace.Used` across all 9 TaskManagers):

| step | metaspace |
|---|---|
| baseline | 3262.6 MB |
| after 3 identical probe jobs | 3317.3 MB (**+54.7 MB, 18.2 MB per job**) |
| after 90 s idle | 3317.3 MB (**0 MB reclaimed**) |

Long-run behaviour, from 15-minute polling — **3,008 samples across 48 distinct TaskManager instances**:

- several TaskManagers climb **monotonically** from ~5 % to 86–92 % of the 768 MB cap and never recover;
- across *all* 48 instances only **9** self-reclamations were ever observed, each just **1.4–4.8 pp**;
- the only thing that returns memory is restarting the TaskManager.

**The decisive contrast — same cluster, same JVM, same GC settings:**

| | ClassesLoaded | ClassesUnloaded | unload ratio |
|---|---|---|---|
| **JobManager** | 359,055 | 241,361 | **67 %** |
| TaskManager #1 | 93,989 | 7,030 | 7.5 % |
| TaskManager #2 | 93,181 | 9,815 | 10.5 % |
| TaskManager #3 | 80,126 | 4,568 | 5.7 % |
| TaskManager #4 | 76,509 | 2,475 | 3.2 % |
| TaskManager #5 | 68,729 | 687 | 1.0 % |
| TaskManager #6 | 25,755 | **0** | **0 %** |
| TaskManager #7 | 26,465 | **0** | **0 %** |
| TaskManager #8 | 23,460 | **0** | **0 %** |

The JobManager in the same cluster unloads two thirds of everything it loads, and was observed reclaiming 150–180 MB at a time **without any restart** (`restartCount=0`, uptime 3 days). The TaskManagers essentially never unload. So this is not GC tuning, not the JVM, and not the G1 configuration — it is user-classloader retention that is specific to the TaskManager side.

### Causes we were able to exclude (so nobody repeats the work)

- **log4j `monitorInterval` (FLINK-32203).** The `log4j-console.properties` actually used by JM/TM contains no `monitorInterval`, and a SIGQUIT thread dump shows **no log4j configuration-watch thread** on the TaskManager. (`monitorInterval=30` exists only in `log4j-cli/session.properties`, i.e. the short-lived client JVM.)
- **JDBC drivers (FLINK-19005).** Reproduces with the Paimon jar alone; no JDBC connector or driver is involved in the query above.
- **Per-job lingering threads.** A TaskManager holding ~570 MB of Metaspace (≈31 retained classloaders at 18.2 MB each) has only **287 JVM threads in total**, with **4** distinct `s3a-*` thread pools and **1** `java-sdk-http-connection-reaper`. Thread counts do **not** scale with the number of retained classloaders, so the mechanism described in #2780 (one unclosed `S3AFileSystem` per job) does not by itself account for what we see here — although an unclosed filesystem would still be a real leak and may explain part of it.
- **"GC simply never ran".** Worth stating explicitly because it is an easy wrong turn: `Status.JVM.GarbageCollector.G1_Old_Generation.Count` is **0 on every TaskManager and on the JobManager**, so that metric cannot be used to argue either way. The JobManager nevertheless unloads 67 % of its classes, which proves reclamation does happen; the TaskManagers just don't benefit from it.

### What we could not do, and why

We cannot name the object that retains the classloader, because the **official Flink image ships a JRE** — there is no `jcmd`, `jmap` or `jstack` in the TaskManager container, so we cannot take a heap dump or run `VM.classloader_stats`.

**If a maintainer wants it, we are happy to run one TaskManager on a JDK-based image and attach a heap dump / `VM.classloader_stats` output, or to test a patch.** Please just say which you'd prefer.

### Why we think Paimon is the right place for this

The retained classloader is the one carrying the Paimon Flink connector jar, and the workload is nothing but "create a Paimon catalog, read one small table, exit". Everything else in the job is stock Flink. #2780 already points at Paimon's S3 filesystem lifecycle (`S3AFileSystem` created but never `close()`d), which would be exactly this class of problem; this report adds a current-version (Paimon 1.2/1.3, Flink 1.20.5) reproducer with a quantified per-job cost and the JM-vs-TM unload contrast, in case that helps narrow it down.

### Impact and current mitigation

In a session cluster used for periodic small Paimon queries this is effectively unbounded: our TaskManagers reach the 768 MB cap and have to be rotated. We currently automate a rolling TaskManager restart at 85 % Metaspace, which takes ~28 s and is transparent to running CDC jobs — but it is obviously a workaround, not a fix.

### Are you willing to submit a PR?

- [X] I'm willing to submit a PR! (happy to help with reproduction, measurements, or testing a fix; we can also run the JDK-image heap dump described above)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the q.sql reproducer and the /opt/flink/bin/sql-client.sh invocation, then run it on a JDK-based TaskManager image to obtain the requested heap dump or VM.classloader_stats output. No Paimon source file or test is named in the report; done means identifying and fixing the retained connector classloader so repeated jobs no longer produce persistent Metaspace growth.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, kubernetes, sql
Domain
backend, data-engineering, distributed-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.