slackapi / slackapi/python-slack-sdk

Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals

Offen
#1,940 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

auto-triage-skip discussion
Vorherrschende Sprache
Python
Sterne
4k
Forks
857
Ø Merge
22 Std. 21 Min.
Gemergte PRs (30 T.)
16

Beschreibung

Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT

Summary

This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.

Wire evidence

(from a run_message_listeners wrapper logging hello and disconnect frames)

  • fresh start: hello reports num_connections=1
  • ~1.5 h later, on a reconnect: 3x disconnect reason=too_many_websockets,
    then hello num_connections=10 — while the process verifiably held ONE
    established TCP connection to Slack the whole time
  • ghost registrations age out at roughly one per 30-45 minutes
  • while num_connections is high, most button clicks never arrive on any
    connection we hold; with a clean pool, every click arrives (tested across
    message sizes 0.5-5 KB — size is irrelevant)
  • reproduced on a SECOND app in the same workspace: first hello after a
    process restart reported num_connections=7 for an app that also runs as a
    single instance
  • observed approximate_connection_time (inside hello.debug_info) is
    consistently 18060 (~5 h), which sets the ghost age-out horizon

Why this is hard to see with the current SDK

  1. The hello envelope (carrying num_connections) never reaches
    message_listenersSocketModeRequest.from_dict requires
    type+envelope_id+payload, so apps cannot observe the most important signal
    without wrapping internals.
  2. disconnect frames (including too_many_websockets) are handled by
    run_message_listeners before the listener loop and only visible at debug
    logging.
  3. A degraded connection still passes is_connected() / ping-pong checks, so
    client-side health monitoring cannot detect the server-side pool state.

Proposals (any subset would help)

  1. Surface hello metadata (num_connections, approximate_connection_time,
    host) and disconnect reasons via a public callback or at INFO logging.
  2. Emit a loud warning when num_connections in hello exceeds a threshold
    (e.g., 4) while the client manages fewer connections — this is direct
    evidence of ghost registrations and imminent interactive-payload loss.
  3. Consider make-before-break reconnects with close-confirmation, or documenting
    that reconnect-heavy operation behind NAT can poison the server-side pool.
  4. Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
    same failure family — this issue is their server-side counterpart, and since
    neither is merged/released (latest release is 3.43.0), apps currently have no
    upstream remedy at all; that raises the priority of surfacing the diagnostics
    from proposals 1-2.

Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne damit, run_message_listeners und SocketModeRequest.from_dict nachzuverfolgen, und konzentriere dich darauf, wie hello-Metadaten und Gründe für Verbindungsabbrüche behandelt werden, bevor Nachrichten-Listener Frames empfangen. Reproduziere das Reconnect-Verhalten, wenn möglich, und definiere anschließend Diagnosen, die die serverseitige Verbindungsanzahl und die Gründe so klar offenlegen, dass vor einem ungesunden Pool gewarnt werden kann.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
backend, networking
Issue-Typ
Bug
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Aktiv
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.