slackapi / slackapi/python-slack-sdk
Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals
Chưa có ai nhận issue này.
- Ngôn ngữ chính
- Python
- Star
- 4k
- Fork
- 857
- Merge trung bình
- 22 giờ 21 phút
- Pull request đã merge (30 ngày)
- 16
Mô tả
Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT
Summary
This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a run_message_listeners wrapper logging hello and disconnect frames)
- fresh start:
helloreportsnum_connections=1 - ~1.5 h later, on a reconnect: 3x
disconnect reason=too_many_websockets,
thenhello num_connections=10— while the process verifiably held ONE
established TCP connection to Slack the whole time - ghost registrations age out at roughly one per 30-45 minutes
- while
num_connectionsis high, most button clicks never arrive on any
connection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant) - reproduced on a SECOND app in the same workspace: first
helloafter a
process restart reportednum_connections=7for an app that also runs as a
single instance - observed
approximate_connection_time(insidehello.debug_info) is
consistently18060(~5 h), which sets the ghost age-out horizon
Why this is hard to see with the current SDK
- The
helloenvelope (carryingnum_connections) never reaches
message_listeners—SocketModeRequest.from_dictrequires
type+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals. disconnectframes (includingtoo_many_websockets) are handled by
run_message_listenersbefore the listener loop and only visible at debug
logging.- A degraded connection still passes
is_connected()/ ping-pong checks, so
client-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
- Surface
hellometadata (num_connections,approximate_connection_time,
host) anddisconnectreasons via a public callback or at INFO logging. - Emit a loud warning when
num_connectionsinhelloexceeds a threshold
(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss. - Consider make-before-break reconnects with close-confirmation, or documenting
that reconnect-heavy operation behind NAT can poison the server-side pool. - Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
same failure family — this issue is their server-side counterpart, and since
neither is merged/released (latest release is 3.43.0), apps currently have no
upstream remedy at all; that raises the priority of surfacing the diagnostics
from proposals 1-2.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Hướng nghiên cứu
Trước tiên, hãy lần theo run_message_listeners và SocketModeRequest.from_dict, tập trung vào cách siêu dữ liệu hello và các lý do ngắt kết nối được xử lý trước khi các listener của thông báo nhận các frame. Nếu có thể, hãy tái hiện hành vi kết nối lại, sau đó xác định các chẩn đoán hiển thị số lượng kết nối phía máy chủ và các lý do đủ rõ ràng để cảnh báo về một pool không lành mạnh.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python
- Lĩnh vực
- backend, networking
- Loại issue
- Lỗi
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức độ hoạt động
- Sôi nổi
- Độ rõ ràng
- Cần làm rõ
- Mức phù hợp với người mới
- 35/100