SAP / SAP/cloud-sdk-java

DefaultHttpDestination.equals() change in 5.34.0 causes MT sidecar 408 timeouts under load

Aperta
#1,268 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Lingua principale
Java
Stelle
41
Fork
33
Merge medio
18h 34m
PR unite (30g)
19

Descrizione

Summary

After upgrading from 5.33.0 to 5.34.0, a Java CAP application that uses the MTX sidecar (@sap/cds-mtxs) for multi-tenancy experiences HTTP 408 Request Timeout responses from the sidecar during tenant subscribe/unsubscribe. The issue is consistently reproducible on a loaded CI environment (Jenkins), though not locally where the sidecar responds fast enough.

Affected versions

  • Broken: 5.34.0
  • Working: 5.33.0

Root cause analysis

PR #1095 ("Fix unexpected connection-pool shut-down") changed DefaultHttpDestination.equals() and hashCode() to now include customHeaderProviders and headerProvidersFromClassLoading in the comparison (previously these were excluded).

DefaultHttpDestination equality is the cache key for the Cloud SDK HTTP client cache (DefaultApacheHttpClient5Cache). MT scenarios attach per-tenant header providers (auth/token headers) to destinations. Before this change, all tenants pointing at the same sidecar URI shared one HttpClient and one connection pool. After this change, each tenant's destination is a distinct cache key → a new HttpClient + connection pool is created per subscribe/unsubscribe call → rapid connection pool churn against a single sidecar process → the sidecar starts returning 408.

The relevant call chain (in com.sap.cds:cds-feature-mt):

ProvisioningService.subscribe()
  → ServiceCallImpl.execute()
  → HttpClientFactory.getHttpClient(destination)     // calls ApacheHttpClient5Accessor
  → sidecar PUT /-/cds/saas-provisioning/tenant/{id}
  ← HTTP 408
  → InternalError("Unexpected return code 408")      // 408 is not in the retry set {500,502,503,504}
  → MtxSidecarDeploymentHandler.onSubscribe() throws
  → HTTP 500 to the subscribe caller

Evidence

Two consecutive CI builds on the same PR branch both failed with the exact same stack:

Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
  at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)
  ...
Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
  at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)

Failing test: com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant (expects HTTP 201, gets 500). All other 13 commits in the 5.33.0→5.34.0 range are dependency bumps that are already overridden by the consuming project's own version pins — the only behaviour-changing commit is #1095.

Steps to reproduce

  1. Run an MTX-sidecar-based CAP Java application's integration tests that subscribe/unsubscribe multiple tenants in rapid succession (e.g. the mtx-local module of cds-services).
  2. Each subscribe/unsubscribe call hits the sidecar via ApacheHttpClient5Accessor.getHttpClient(destination) where the destination carries tenant-specific header providers.
  3. With 5.34.0, a new HttpClient + connection pool is allocated per tenant on every call → pool exhaustion / timeout after several tenants → 408 from the sidecar.
  4. With 5.33.0, all same-URI destinations share one HttpClient and pool → no exhaustion → sidecar responds 200/202.

Suggested fix

Options:

  • On the cloud-sdk side: consider whether the HTTP-client cache should key on something coarser (e.g. URI only, or URI + a stable identity of the header providers) rather than full header provider equality, especially for the per-request-dynamic providers used in MT scenarios.
  • As a workaround: the consuming application can pin cloud.sdk.version=5.33.0 until this is resolved.

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Direzione di ricerca

Inizia confrontando le modifiche a DefaultHttpDestination.equals() e hashCode() in PR #1095 con il comportamento di DefaultApacheHttpClient5Cache. Riproduci il malfunzionamento tramite com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant e traccia la destination passata a ApacheHttpClient5Accessor.getHttpClient(). Il lavoro è completato quando lo scenario MT di subscribe/unsubscribe non produce più risposte 408 o HTTP 500 sotto carico.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
backend, networking
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
48/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.