splunk / splunk/splunk-sdk-python

connect() fails on FIPS-enabled Splunk 10.4.x: unguarded ctx.set_groups() raises ssl.SSLError (_ssl.c:4981) before any connection is attempted

Ouverte Adaptée aux débutants
#835 1 commentaire 1 réaction 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

Langage dominant
Python
Étoiles
743
Forks
387
Merge moyen
42 min
PR mergées (30 j)
4

Description

Summary

splunklib/binding.py connect() (the verify=False path — the splunklib default) configures an explicit TLS key-exchange group list that includes ML-KEM hybrid groups:

if hasattr(ctx, "set_groups"):
    ctx.set_groups(  # pyright: ignore[reportUnknownMemberType, reportAttributeAccessIssue]
        "X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
        + "MLKEM512:MLKEM768:MLKEM1024:"
        + "X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072"
    )

(binding.py lines 1744-1758 in splunk-sdk 3.0.0; identical on develop today.)

As the adjacent comment notes, SSLContext.set_groups is a Python 3.15 API that Splunk backports into its patched 3.9/3.13 runtimes — so on Splunk builds that carry the backport, the hasattr guard passes. But there is no exception handling: on a FIPS-enabled Splunk deployment, the OpenSSL FIPS provider has no validated ML-KEM, SSL_CTX_set1_groups_list rejects the group list, and every client.connect() (and any app-level code using splunklib against splunkd) dies before a single packet is sent:

File ".../lib/splunklib/binding.py", line 1754, in connect
    ctx.set_groups(
ssl.SSLError: unknown error (_ssl.c:4981)

This breaks any Splunk app that bundles splunk-sdk 3.0.0 on a FIPS-enabled Splunk 10.4.x search head — in our case, all REST handlers of a widely deployed app became inoperable after the customer-side platform upgrade, surfacing as HTTP 500 on every request.

Environment / verified matrix

Splunk Runtime OpenSSL hasattr(ctx, "set_groups") ctx.set_groups(<SDK list>)
10.2.4 (1526e5e5df42), FIPS bundled Python 3.9.25 and 3.13.11 3.0.19 False n/a — SDK skips the call, everything works
10.4.2 (33c3bf42cd73), FIPS bundled Python 3.13.11 (what python3 resolves to) 3.5.7 True SSLError(0, 'unknown error (_ssl.c:4981)')

Per-group-name bisection on 10.4.2 FIPS (splunk cmd python3, ssl._create_unverified_context()):

  • FAIL: X25519MLKEM768, SecP256r1MLKEM768, SecP384r1MLKEM1024, MLKEM512, MLKEM768, MLKEM1024 — each raises SSLError(0, 'unknown error (_ssl.c:4981)')
  • OK: X25519, secp256r1, X448, secp384r1, secp521r1, ffdhe2048, ffdhe3072

So the failure needs no app code at all to reproduce — the FIPS provider simply rejects every ML-KEM group name in the hardcoded list.

Minimal reproduction (on a FIPS-enabled Splunk 10.4.2)

$SPLUNK_HOME/bin/splunk cmd python3 - <<'EOF'
import ssl
ctx = ssl._create_unverified_context()
print("has set_groups:", hasattr(ctx, "set_groups"))   # True on 10.4.2
ctx.set_groups("X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
               "MLKEM512:MLKEM768:MLKEM1024:"
               "X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072")
EOF
# → ssl.SSLError: unknown error (_ssl.c:4981)

Any splunklib.client.connect(...) on such a host then fails identically inside binding.py connect().

Suggested fix

Wrap the call so a rejected group list degrades to the context's default groups — the exact behaviour every pre-backport Splunk runs today:

if hasattr(ctx, "set_groups"):
    try:
        ctx.set_groups(
            "X25519MLKEM768:SecP256r1MLKEM768:SecP384r1MLKEM1024:"
            + "MLKEM512:MLKEM768:MLKEM1024:"
            + "X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072"
        )
    except ssl.SSLError:
        # FIPS providers (and OpenSSL builds without ML-KEM) reject the
        # ML-KEM group names — fall back to the default groups rather
        # than failing every connection.
        pass

Optionally retry with the classical-only tail (X25519:secp256r1:X448:secp384r1:secp521r1:ffdhe2048:ffdhe3072 — all accepted by the FIPS provider in our bisection) before falling back entirely, but a plain fallback already restores full functionality. Happy to submit a PR for either variant.

Workaround we ship meanwhile

Our app's build pipeline patches the bundled binding.py to add exactly the try/except above. Validated: the patched build fully restores functionality on the FIPS-enabled Splunk 10.4.2 reproduction host (same host, same runtime — all REST handlers operational again), confirming the fallback-to-default-groups approach is sufficient.


AI disclosure

This issue was investigated, reproduced, and written with the assistance of an AI coding agent (Claude Code by Anthropic), operated and reviewed by Guilhem Marchand (TrackMe Limited). All reproduction outputs quoted above were produced on real Splunk hosts and human-verified before filing.

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez dans splunklib/binding.py, au niveau de connect(), lignes 1744-1758, et exécutez la reproduction minimale de set_groups sur un runtime Splunk 10.4.2 avec FIPS activé. Gérez l’échec de la configuration du groupe afin que les connexions reviennent aux valeurs par défaut du contexte, puis vérifiez que client.connect() et les handlers REST concernés continuent sans l’erreur SSL.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
security
Type d'issue
Bug
Difficulté
2/5
Temps estimé
1-3 heures
Activité
Active
Clarté
Clairement spécifiée
Accessibilité débutants
78/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.