DefaultHttpDestination.equals() change in 5.34.0 causes MT sidecar 408 timeouts under load
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- java
- Área
- backend, networking
Línea de trabajo
Empieza comparando los cambios en DefaultHttpDestination.equals() y hashCode() de PR #1095 con el comportamiento de DefaultApacheHttpClient5Cache. Reproduce el fallo mediante com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant y sigue la destination que se pasa a ApacheHttpClient5Accessor.getHttpClient(). Se considera terminado cuando el escenario de subscribe/unsubscribe de MT ya no produce respuestas 408 ni HTTP 500 bajo carga.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
After upgrading from 5.33.0 to 5.34.0, a Java CAP application that uses the MTX sidecar (@sap/cds-mtxs) for multi-tenancy experiences HTTP 408 Request Timeout responses from the sidecar during tenant subscribe/unsubscribe. The issue is consistently reproducible on a loaded CI environment (Jenkins), though not locally where the sidecar responds fast enough.
Affected versions
- Broken:
5.34.0 - Working:
5.33.0
Root cause analysis
PR #1095 ("Fix unexpected connection-pool shut-down") changed DefaultHttpDestination.equals() and hashCode() to now include customHeaderProviders and headerProvidersFromClassLoading in the comparison (previously these were excluded).
DefaultHttpDestination equality is the cache key for the Cloud SDK HTTP client cache (DefaultApacheHttpClient5Cache). MT scenarios attach per-tenant header providers (auth/token headers) to destinations. Before this change, all tenants pointing at the same sidecar URI shared one HttpClient and one connection pool. After this change, each tenant's destination is a distinct cache key → a new HttpClient + connection pool is created per subscribe/unsubscribe call → rapid connection pool churn against a single sidecar process → the sidecar starts returning 408.
The relevant call chain (in com.sap.cds:cds-feature-mt):
ProvisioningService.subscribe()
→ ServiceCallImpl.execute()
→ HttpClientFactory.getHttpClient(destination) // calls ApacheHttpClient5Accessor
→ sidecar PUT /-/cds/saas-provisioning/tenant/{id}
← HTTP 408
→ InternalError("Unexpected return code 408") // 408 is not in the retry set {500,502,503,504}
→ MtxSidecarDeploymentHandler.onSubscribe() throws
→ HTTP 500 to the subscribe caller
Evidence
Two consecutive CI builds on the same PR branch both failed with the exact same stack:
Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)
...
Caused by: com.sap.cds.feature.mt.lib.subscription.exceptions.InternalError: Unexpected return code 408
at com.sap.cds.feature.mt.lib.subscription.ProvisioningService.lambda$new$1(ProvisioningService.java:80)
Failing test: com.sap.mtx.multitenancy.SubscribeAndUnsubscribeTest.onBoardAndOffBoardNewTenant (expects HTTP 201, gets 500). All other 13 commits in the 5.33.0→5.34.0 range are dependency bumps that are already overridden by the consuming project's own version pins — the only behaviour-changing commit is #1095.
Steps to reproduce
- Run an MTX-sidecar-based CAP Java application's integration tests that subscribe/unsubscribe multiple tenants in rapid succession (e.g. the
mtx-localmodule ofcds-services). - Each subscribe/unsubscribe call hits the sidecar via
ApacheHttpClient5Accessor.getHttpClient(destination)where the destination carries tenant-specific header providers. - With
5.34.0, a new HttpClient + connection pool is allocated per tenant on every call → pool exhaustion / timeout after several tenants → 408 from the sidecar. - With
5.33.0, all same-URI destinations share one HttpClient and pool → no exhaustion → sidecar responds 200/202.
Suggested fix
Options:
- On the cloud-sdk side: consider whether the HTTP-client cache should key on something coarser (e.g. URI only, or URI + a stable identity of the header providers) rather than full header provider equality, especially for the per-request-dynamic providers used in MT scenarios.
- As a workaround: the consuming application can pin
cloud.sdk.version=5.33.0until this is resolved.
- Lenguaje dominante
- Java
- Estrellas
- 41
- Forks
- 33
- Merge medio
- 18 h 34 min
- PR fusionados (30 d)
- 19
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de SAP/cloud-sdk-java
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
SAP/cloud-sdk-java#1280 · 1 comentario ·
-
bug
SAP/cloud-sdk-java#1270 · 3 comentarios · 1 asignado ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
SAP/cloud-sdk-java#1250 ·
-
ZeroTrustIdentityService does not configure svidPicker, causing non-deterministic SVID selection Abiertobug
Dificultad 3/5 1-2 días Aptitud para principiantes 70/100
SAP/cloud-sdk-java#1243 · 2 comentarios ·
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
SAP/cloud-sdk-java#1226 · 1 comentario ·
Todos los issues de SAP/cloud-sdk-java
Issues similares
-
Bug Java Platform: Java
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
getsentry/sentry-java#6138 · 1 comentario ·
-
bug needs triage p2
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
GoogleCloudPlatform/DataflowTemplates#4273 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
-
bug needs triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100