slackapi / slackapi/python-slack-sdk

Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals

Ouverte
#1,940 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

auto-triage-skip discussion
Langage dominant
Python
Étoiles
4k
Forks
857
Merge moyen
22 h 21 min
PR mergées (30 j)
16

Description

Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT

Summary

This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.

Wire evidence

(from a run_message_listeners wrapper logging hello and disconnect frames)

  • fresh start: hello reports num_connections=1
  • ~1.5 h later, on a reconnect: 3x disconnect reason=too_many_websockets,
    then hello num_connections=10 — while the process verifiably held ONE
    established TCP connection to Slack the whole time
  • ghost registrations age out at roughly one per 30-45 minutes
  • while num_connections is high, most button clicks never arrive on any
    connection we hold; with a clean pool, every click arrives (tested across
    message sizes 0.5-5 KB — size is irrelevant)
  • reproduced on a SECOND app in the same workspace: first hello after a
    process restart reported num_connections=7 for an app that also runs as a
    single instance
  • observed approximate_connection_time (inside hello.debug_info) is
    consistently 18060 (~5 h), which sets the ghost age-out horizon

Why this is hard to see with the current SDK

  1. The hello envelope (carrying num_connections) never reaches
    message_listenersSocketModeRequest.from_dict requires
    type+envelope_id+payload, so apps cannot observe the most important signal
    without wrapping internals.
  2. disconnect frames (including too_many_websockets) are handled by
    run_message_listeners before the listener loop and only visible at debug
    logging.
  3. A degraded connection still passes is_connected() / ping-pong checks, so
    client-side health monitoring cannot detect the server-side pool state.

Proposals (any subset would help)

  1. Surface hello metadata (num_connections, approximate_connection_time,
    host) and disconnect reasons via a public callback or at INFO logging.
  2. Emit a loud warning when num_connections in hello exceeds a threshold
    (e.g., 4) while the client manages fewer connections — this is direct
    evidence of ghost registrations and imminent interactive-payload loss.
  3. Consider make-before-break reconnects with close-confirmation, or documenting
    that reconnect-heavy operation behind NAT can poison the server-side pool.
  4. Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
    same failure family — this issue is their server-side counterpart, and since
    neither is merged/released (latest release is 3.43.0), apps currently have no
    upstream remedy at all; that raises the priority of surfacing the diagnostics
    from proposals 1-2.

Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.

Guide de contribution

Ouvrir le guide de contribution

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Piste de recherche

Commencez par retracer run_message_listeners et SocketModeRequest.from_dict, en vous concentrant sur la manière dont les métadonnées de hello et les raisons de déconnexion sont gérées avant que les listeners de messages ne reçoivent des frames. Reproduisez le comportement de reconnexion si possible, puis définissez des diagnostics qui exposent le nombre de connexions côté serveur et les raisons suffisamment clairement pour avertir qu’un pool est défaillant.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
python
Domaine
backend, networking
Type d'issue
Bug
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
Active
Clarté
À clarifier
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.