apache / apache/openserverless

Installation fails on Ubuntu 24.04 + kernel 6.14 - JVM cgroup v2 compatibility issue

オープン
#193 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

主要言語
Python
スター
576
フォーク
29
平均マージ
50分
マージ済み PR(30日)
13

説明

Ⅰ. Issue Description

OpenServerless installation fails on Ubuntu 24.04 with kernel 6.14 due to JVM cgroup v2 compatibility issue. Kafka and Controller pods crash with NullPointerException in CgroupV2Subsystem.getInstance.

Ⅱ. Describe what happened

Kafka pod crashes immediately on startup with the following exception:

Exception in thread "main" java.lang.reflect.InvocationTargetException
    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
    at java.base/jdk.internal.reflect.NativeMethodAccessorImpl.invoke(Unknown Source)
    at java.base/jdk.internal.reflect.DelegatingMethodAccessorImpl.invoke(Unknown Source)
    at java.base/java.lang.reflect.Method.invoke(Unknown Source)
    at java.instrument/sun.instrument.InstrumentationImpl.loadClassAndStartAgent(Unknown Source)
    at java.instrument/sun.instrument.InstrumentationImpl.loadClassAndCallPremain(Unknown Source)
Caused by: java.lang.NullPointerException
    at java.base/jdk.internal.platform.cgroupv2.CgroupV2Subsystem.getInstance(Unknown Source)
    at java.base/jdk.internal.platform.CgroupSubsystemFactory.create(Unknown Source)
    at java.base/jdk.internal.platform.CgroupMetrics.getInstance(Unknown Source)
    at java.base/jdk.internal.platform.SystemMetrics.instance(Unknown Source)
    at java.base/jdk.internal.platform.Metrics.systemMetrics(Unknown Source)
    at java.base/jdk.internal.platform.Container.metrics(Unknown Source)
    at jdk.management/com.sun.management.internal.OperatingSystemImpl.<init>(Unknown Source)
    at jdk.management/com.sun.management.internal.PlatformMBeanProviderImpl.getOperatingSystemMXBean(Unknown Source)
    at jdk.management/com.sun.management.internal.PlatformMBeanProviderImpl$3.nameToMBeanMap(Unknown Source)
    at java.management/sun.management.spi.PlatformMBeanProvider$PlatformComponent.getMBeans(Unknown Source)
    at java.management/java.lang.management.ManagementFactory.getPlatformMXBean(Unknown Source)
    at java.management/java.lang.management.ManagementFactory.getOperatingSystemMXBean(Unknown Source)
    at io.prometheus.jmx.shaded.io.prometheus.client.hotspot.StandardExports.<init>(StandardExports.java:43)
    at io.prometheus.jmx.shaded.io.prometheus.client.hotspot.DefaultExports.register(DefaultExports.java:37)
    at io.prometheus.jmx.shaded.io.prometheus.client.hotspot.DefaultExports.initialize(DefaultExports.java:28)
    at io.prometheus.jmx.JavaAgent.premain(JavaAgent.java:30)
    ... 6 more
*** java.lang.instrument ASSERTION FAILED ***: "result" with message agent load/premain call failed at src/java.instrument/share/native/libinstrument/JPLISAgent.c line: 422
FATAL ERROR in native method: processing of -javaagent failed, processJavaStart failed

The Controller pod never starts - installation hangs indefinitely waiting for pod/controller-0 to become available.

Ⅲ. Describe what you expected to happen

Installation should complete successfully with all pods running and healthy, allowing user creation and login.

Ⅳ. How to reproduce it (as minimally and precisely as possible)
  1. Install Ubuntu 24.04 on hardware with NVIDIA kernel 6.14
  2. Configure ops with minimal settings
  3. Run installation
  4. Installation fails with timeout waiting for controller-0
Ⅴ. Anything else we need to know?

Root Cause Analysis:

This is a known issue with JVM versions prior to JDK 21 running on Linux kernel 6.12+ with cgroup v2. The problem occurs because:

  1. Ubuntu 24.04 uses kernel 6.x + systemd 256+ which enforces cgroup v2
  2. The memory cgroup controller is missing from /proc/cgroups:
   $ cat /proc/cgroups
   #subsys_name    hierarchy    num_cgroups    enabled
cpu    0    860    1
cpuacct    0    860    1
blkio    0    860    1
devices    0    860    1
freezer    0    860    1
net_cls    0    860    1
perf_event    0    860    1
net_prio    0    860    1
hugetlb    0    860    1
pids    0    860    1
rdma    0    860    1
misc    0    860    1
dmem    0    860    1

Note: No "memory" controller present
3. JVM's CgroupV2Subsystem.getInstance() expects the memory controller and crashes with NPE when it's missing

Temporary Workaround:

Force cgroup v1 by adding kernel boot parameter:

# In /etc/default/grub:
GRUB_CMDLINE_LINUX="systemd.unified_cgroup_hierarchy=0"
# Then: sudo update-grub && sudo reboot

Permanent Solution:

Update Docker images to use JDK 21+ or apply JVM patches for cgroup v2 compatibility. Similar issues have been fixed in:

Ⅵ. Environment:
  • K8S Runtime and version: k3s (installed by ops)
  • OPS CLI version: 0.1.0-2409121919.dev
  • OS: Ubuntu 24.04 LTS
  • Kernel: 6.14.0-1015-nvidia
  • Hardware: Dell Pro Max GB10 (NVIDIA GB10 Grace CPU + Blackwell GPU)
  • Java version on host: OpenJDK 1.8.0_472
  • Cgroup version: v2 (enforced by kernel 6.14)

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

Kafka および Controller Pod に使用されている Docker イメージとインストールのエントリーポイントを特定し、続いてそれらにバンドルされている JVM バージョンを調べます。cgroup v2 を使用する Ubuntu 24.04 で失敗を再現し、選択したイメージの変更によって controller-0 とその他の Pod が正常な状態でインストールを完了できることを検証します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
docker, java, kubernetes, linux
領域
cloud, devops, distributed-systems, infrastructure
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。