splunk / splunk/splunk-sdk-python
connect() fails on FIPS-enabled Splunk 10.4.x: unguarded ctx.set_groups() raises ssl.SSLError (_ssl.c:4981) before any connection is attempted
Personne n'a encore pris cette issue.
- Langage dominant
- Python
- Étoiles
- 743
- Forks
- 387
- Merge moyen
- 42 min
- PR mergées (30 j)
- 4
Description
Summary
splunklib/binding.py connect() (the verify=False path — the splunklib default) configures an explicit TLS key-exchange group list that includes ML-KEM hybrid groups:
if hasattr(ctx, "set_groups"):
ctx.set_groups( # pyright: ignore[reportUnknownMemberType, reportAttributeAccessIssue]
"X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
+ "MLKEM512:MLKEM768:MLKEM1024:"
+ "X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072"
)
(binding.py lines 1744-1758 in splunk-sdk 3.0.0; identical on develop today.)
As the adjacent comment notes, SSLContext.set_groups is a Python 3.15 API that Splunk backports into its patched 3.9/3.13 runtimes — so on Splunk builds that carry the backport, the hasattr guard passes. But there is no exception handling: on a FIPS-enabled Splunk deployment, the OpenSSL FIPS provider has no validated ML-KEM, SSL_CTX_set1_groups_list rejects the group list, and every client.connect() (and any app-level code using splunklib against splunkd) dies before a single packet is sent:
File ".../lib/splunklib/binding.py", line 1754, in connect
ctx.set_groups(
ssl.SSLError: unknown error (_ssl.c:4981)
This breaks any Splunk app that bundles splunk-sdk 3.0.0 on a FIPS-enabled Splunk 10.4.x search head — in our case, all REST handlers of a widely deployed app became inoperable after the customer-side platform upgrade, surfacing as HTTP 500 on every request.
Environment / verified matrix
| Splunk | Runtime | OpenSSL | hasattr(ctx, "set_groups") |
ctx.set_groups(<SDK list>) |
|---|---|---|---|---|
10.2.4 (1526e5e5df42), FIPS |
bundled Python 3.9.25 and 3.13.11 | 3.0.19 | False |
n/a — SDK skips the call, everything works |
10.4.2 (33c3bf42cd73), FIPS |
bundled Python 3.13.11 (what python3 resolves to) |
3.5.7 | True |
SSLError(0, 'unknown error (_ssl.c:4981)') |
Per-group-name bisection on 10.4.2 FIPS (splunk cmd python3, ssl._create_unverified_context()):
- FAIL:
X25519MLKEM768,SecP256r1MLKEM768,SecP384r1MLKEM1024,MLKEM512,MLKEM768,MLKEM1024— each raisesSSLError(0, 'unknown error (_ssl.c:4981)') - OK:
X25519,secp256r1,X448,secp384r1,secp521r1,ffdhe2048,ffdhe3072
So the failure needs no app code at all to reproduce — the FIPS provider simply rejects every ML-KEM group name in the hardcoded list.
Minimal reproduction (on a FIPS-enabled Splunk 10.4.2)
$SPLUNK_HOME/bin/splunk cmd python3 - <<'EOF'
import ssl
ctx = ssl._create_unverified_context()
print("has set_groups:", hasattr(ctx, "set_groups")) # True on 10.4.2
ctx.set_groups("X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
"MLKEM512:MLKEM768:MLKEM1024:"
"X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072")
EOF
# → ssl.SSLError: unknown error (_ssl.c:4981)
Any splunklib.client.connect(...) on such a host then fails identically inside binding.py connect().
Suggested fix
Wrap the call so a rejected group list degrades to the context's default groups — the exact behaviour every pre-backport Splunk runs today:
if hasattr(ctx, "set_groups"):
try:
ctx.set_groups(
"X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
+ "MLKEM512:MLKEM768:MLKEM1024:"
+ "X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072"
)
except ssl.SSLError:
# FIPS providers (and OpenSSL builds without ML-KEM) reject the
# ML-KEM group names — fall back to the default groups rather
# than failing every connection.
pass
Optionally retry with the classical-only tail (X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072 — all accepted by the FIPS provider in our bisection) before falling back entirely, but a plain fallback already restores full functionality. Happy to submit a PR for either variant.
Workaround we ship meanwhile
Our app's build pipeline patches the bundled binding.py to add exactly the try/except above. Validated: the patched build fully restores functionality on the FIPS-enabled Splunk 10.4.2 reproduction host (same host, same runtime — all REST handlers operational again), confirming the fallback-to-default-groups approach is sufficient.
AI disclosure
This issue was investigated, reproduced, and written with the assistance of an AI coding agent (Claude Code by Anthropic), operated and reviewed by Guilhem Marchand (TrackMe Limited). All reproduction outputs quoted above were produced on real Splunk hosts and human-verified before filing.
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez dans splunklib/binding.py, au niveau de connect(), lignes 1744-1758, et exécutez la reproduction minimale de set_groups sur un runtime Splunk 10.4.2 avec FIPS activé. Gérez l’échec de la configuration du groupe afin que les connexions reviennent aux valeurs par défaut du contexte, puis vérifiez que client.connect() et les handlers REST concernés continuent sans l’erreur SSL.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- security
- Type d'issue
- Bug
- Difficulté
- 2/5
- Temps estimé
- 1-3 heures
- Activité
- Active
- Clarté
- Clairement spécifiée
- Accessibilité débutants
- 78/100