slackapi / slackapi/python-slack-sdk
Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals
Personne n'a encore pris cette issue.
- Langage dominant
- Python
- Étoiles
- 4k
- Forks
- 857
- Merge moyen
- 22 h 21 min
- PR mergées (30 j)
- 16
Description
Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT
Summary
This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a run_message_listeners wrapper logging hello and disconnect frames)
- fresh start:
helloreportsnum_connections=1 - ~1.5 h later, on a reconnect: 3x
disconnect reason=too_many_websockets,
thenhello num_connections=10— while the process verifiably held ONE
established TCP connection to Slack the whole time - ghost registrations age out at roughly one per 30-45 minutes
- while
num_connectionsis high, most button clicks never arrive on any
connection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant) - reproduced on a SECOND app in the same workspace: first
helloafter a
process restart reportednum_connections=7for an app that also runs as a
single instance - observed
approximate_connection_time(insidehello.debug_info) is
consistently18060(~5 h), which sets the ghost age-out horizon
Why this is hard to see with the current SDK
- The
helloenvelope (carryingnum_connections) never reaches
message_listeners—SocketModeRequest.from_dictrequires
type+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals. disconnectframes (includingtoo_many_websockets) are handled by
run_message_listenersbefore the listener loop and only visible at debug
logging.- A degraded connection still passes
is_connected()/ ping-pong checks, so
client-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
- Surface
hellometadata (num_connections,approximate_connection_time,
host) anddisconnectreasons via a public callback or at INFO logging. - Emit a loud warning when
num_connectionsinhelloexceeds a threshold
(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss. - Consider make-before-break reconnects with close-confirmation, or documenting
that reconnect-heavy operation behind NAT can poison the server-side pool. - Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
same failure family — this issue is their server-side counterpart, and since
neither is merged/released (latest release is 3.43.0), apps currently have no
upstream remedy at all; that raises the priority of surfacing the diagnostics
from proposals 1-2.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez par retracer run_message_listeners et SocketModeRequest.from_dict, en vous concentrant sur la manière dont les métadonnées de hello et les raisons de déconnexion sont gérées avant que les listeners de messages ne reçoivent des frames. Reproduisez le comportement de reconnexion si possible, puis définissez des diagnostics qui exposent le nombre de connexions côté serveur et les raisons suffisamment clairement pour avertir qu’un pool est défaillant.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- backend, networking
- Type d'issue
- Bug
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Active
- Clarté
- À clarifier
- Accessibilité débutants
- 35/100