operator-framework / operator-framework/java-operator-sdk

Custom bootstrap and CR lifecycle management

Đang mở
#2,884 9 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Ngôn ngữ chính
Java
Star
944
Fork
242
Merge trung bình
1 ngày 4 giờ
Pull request đã merge (30 ngày)
43

Mô tả

This is mostly a reality check if our custom bootstrap is using JOSDK as intended or not. Depending on the feedback we can close this issue and create more specific follow-up issues if needed. I just wanted to describe the context once.

// configure operator
final var customMetrics = new CustomResourceMetrics(config, meterRegistry);
final var eventSender = new EventSender(config, client);
final var reconciler = new HiveMQPlatformReconciler(config, customResourceMetrics, client, eventSender);

final var dependentResourceFactory = new HiveMQPlatformOperatorDependentResourceFactory<>(config, eventSender);
final var metrics = MicrometerMetrics.newPerResourceCollectingMicrometerMetricsBuilder(meterRegistry).build();
final var operator = new Operator(override -> override //
        .withConcurrentReconciliationThreads(config.getConcurrentReconciliationThreads())
        .withConcurrentWorkflowExecutorThreads(config.getConcurrentWorkflowThreads())
        .withReconciliationTerminationTimeout(config.getReconciliationTerminationTimeout())
        .withCacheSyncTimeout(config.getCacheSyncTimeout())
        .withKubernetesClient(client)
        .withCloseClientOnStop(closeClient)
        .withDependentResourceFactory(dependentResourceFactory)
        .withMetrics(metrics)
        .withUseSSAToPatchPrimaryResource(useSSA)
        .withSSABasedCreateUpdateMatchForDependentResources(useSSA));
final var reconcilerConfig = operator.getConfigurationService().getConfigurationFor(reconciler); // (1)
final var overrideConfig = ControllerConfigurationOverrider.override(reconcilerConfig);
if (operatorNamespaces.equals(Constants.WATCH_CURRENT_NAMESPACE)) {
    overrideConfig.watchingOnlyCurrentNamespace();
} else if (operatorNamespaces.equals(Constants.WATCH_ALL_NAMESPACES)) {
    overrideConfig.watchingAllNamespaces();
} else {
    final var namespaces = Arrays.stream(operatorNamespaces.split(",")).collect(Collectors.toSet());
    overrideConfig.settingNamespaces(namespaces);
}
if (!operatorSelector.isBlank()) {
    overrideConfig.withLabelSelector(operatorSelector);
}
overrideConfig.withOnAddFilter(platform -> { // (3)
    customResourceMetrics.register(platform);
    return true;
});
final var controller = (Controller<HiveMQPlatform>) operator.register(reconciler, overrideConfig.build());
customResourceMetrics.setCache(controller.getEventSourceManager().getControllerEventSource()); // (2)
  • HiveMQPlatformOperatorDependentResourceFactory implements DependentResourceFactory (it's the only place where we really had to replace Quarkus Arc, otherwise we don't need a CDI).
  • closeClient is true in production and only set to false in tests to not close an injected K8s client (that is still used in the test).
  • useSSA is true in production and only set to false in tests with the K8s mockserver.
  1. Are we creating the configuration and overrides as intended? I tried to reverse engineer how Quarkus and the LocallyRunOperatorExtension are configuring JOSDK, but we still get a warning on the operator.getConfigurationService().getConfigurationFor(reconciler) invocation:
    12:29:37.423 [main] WARN  Default ConfigurationService implementation - Configuration for reconciler
    'hivemq-controller' was not found. Known reconcilers: None.
    
    It feels wrong to get that warning, but I found no other way to create a ControllerConfigurationOverrider.
  2. How to properly implement custom metrics on all custom resources in an efficient way? For example, in CustomResourceMetrics we have a global Gauge for each state of our operator state machine. To calculate the values we count all custom resources that are in that state:
    cache.list()
        .map(CustomResource::getStatus)
        .filter(Objects::nonNull)
        .filter(status -> status.getState().equals(state))
        .count();
    
    That cache is the ControllerEventSource that we set via customResourceMetrics.setCache(controller.getEventSourceManager().getControllerEventSource()). This feels a bit illegal but works very fine for us.
    In the previous implementation we kept a shadow copy of all custom resources in a CHM, but these were the cloned instances from the reconciliation loop. This impacted the GC and needed constant updates of the cached instances to retrieve the current state. Using the original instances from the cache solved this problem very elegantly for us, including the lifecycle management.
  3. The same problem hits us now again with a new Gauge for an aggregated health status metric per custom resource. We need to register this Gauge once when the CR is added, de-register it when it's removed and in between have access to the current CustomResource::getStatus to get the health state. So we're back to a Map<String, GaugeHolder> with <namespace>-<name> as key.
    For the Gauge creation I'm using overrideConfig.withOnAddFilter(), which again feels very wrong, but seems to work fine. The Gauge removal is done in the Cleaner::cleanup method of our reconciler.
    Is there a better way for the lifecycle control and where to put that Gauge? It feels like we cannot put this into the custom resource itself, because of the cloning (the Gauge needs to be a singleton) and to update the correct instance. This is how it looks like in the CustomResourceMetrics right now:
    private final @NotNull Map<String, GaugeHolder> healthMetricGauges = new ConcurrentHashMap<>();
    
    // called from the bootstrap
    public void register(final @NotNull HiveMQPlatform platform) {
        healthMetricGauges.put(getKey(platform), new GaugeHolder(platform));
    }
    
    // called from Cleaner::cleanup
    public void deregister(final @NotNull HiveMQPlatform platform) {
        if (healthMetricGauges.get(getKey(platform)) instanceof GaugeHolder gaugeHolder) {
            meterRegistry.remove(gaugeHolder.gauge);
        }
    }
    
    // called from the main reconciler when the health state has changed
    public void updateHealthMetric(final @NotNull HiveMQPlatform platform) {
        if (healthMetricGauges.get(getKey(platform)) instanceof GaugeHolder gaugeHolder) {
            gaugeHolder.value.set(platform.getStatus().getHealthStatus().getMetricsValue());
        }
    }
    
    private static @NotNull String getKey(final @NotNull HiveMQPlatform platform) {
        return platform.getMetadata().getNamespace() + "-" + platform.getMetadata().getName();
    }
    
    private class GaugeHolder {
    
        private final @NotNull AtomicInteger value = new AtomicInteger(HealthStatus.UNKNOWN.getMetricsValue());
        private final @NotNull Gauge gauge;
    
        public GaugeHolder(final @NotNull HiveMQPlatform platform) {
            this.gauge = Gauge.builder("hivemq.platform.health.system.current", value::get)
                    .tag("namespace", platform.getMetadata().getNamespace())
                    .tag("name", platform.getMetadata().getName())
                    .description("Custom resource health status")
                    .register(meterRegistry);
        }
    }
    

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu với mã bootstrap tùy chỉnh, đặc biệt là ControllerConfigurationOverrider và luồng đăng ký operator. Đọc CustomResourceMetrics và Cleaner::cleanup để hiểu vòng đời đăng ký, cập nhật và hủy đăng ký hiện tại. Công việc được xem là hoàn tất khi một cách tiếp cận đã được thống nhất và được JOSDK hỗ trợ cho cấu hình, các metric dựa trên cache và quyền sở hữu gauge theo từng resource được ghi lại hoặc triển khai.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
java, kubernetes
Lĩnh vực
devops, infrastructure, observability
Loại issue
Tái cấu trúc
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.