Socket Mode: reconnects behind NAT leak server-side connection registrations → too_many_websockets cap → silent loss of interactive payloads; mitigation proposals
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- backend, networking
Hướng nghiên cứu
Trước tiên, hãy lần theo run_message_listeners và SocketModeRequest.from_dict, tập trung vào cách siêu dữ liệu hello và các lý do ngắt kết nối được xử lý trước khi các listener của thông báo nhận các frame. Nếu có thể, hãy tái hiện hành vi kết nối lại, sau đó xác định các chẩn đoán hiển thị số lượng kết nối phía máy chủ và các lý do đủ rõ ràng để cảnh báo về một pool không lành mạnh.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Slack SDK version: slack-sdk 3.43.0, slack-bolt 1.29.0
Python: 3.13 (aiohttp Socket Mode client)
OS/platform: Linux container (Docker) on macOS host, behind NAT
Summary
This is part bug report, part mitigation proposal, backed by wire-level data.
When SocketModeClient reconnects while the network path is degraded (half-open
TCP: the close frame never reaches Slack), the old connection remains registered
server-side for an extended period. Repeated reconnects therefore accumulate
"ghost" registrations up to Slack's 10-connection cap (disconnect: too_many_websockets). Slack then delivers envelopes across all registered
connections, so most interactive (block_actions) payloads — which unlike
events are not retried on non-ack — are silently lost. From the app's
perspective the client looks perfectly healthy: ping/pong fine, events flowing.
Wire evidence
(from a run_message_listeners wrapper logging hello and disconnect frames)
- fresh start:
helloreportsnum_connections=1 - ~1.5 h later, on a reconnect: 3x
disconnect reason=too_many_websockets,
thenhello num_connections=10— while the process verifiably held ONE
established TCP connection to Slack the whole time - ghost registrations age out at roughly one per 30-45 minutes
- while
num_connectionsis high, most button clicks never arrive on any
connection we hold; with a clean pool, every click arrives (tested across
message sizes 0.5-5 KB — size is irrelevant) - reproduced on a SECOND app in the same workspace: first
helloafter a
process restart reportednum_connections=7for an app that also runs as a
single instance - observed
approximate_connection_time(insidehello.debug_info) is
consistently18060(~5 h), which sets the ghost age-out horizon
Why this is hard to see with the current SDK
- The
helloenvelope (carryingnum_connections) never reaches
message_listeners—SocketModeRequest.from_dictrequires
type+envelope_id+payload, so apps cannot observe the most important signal
without wrapping internals. disconnectframes (includingtoo_many_websockets) are handled by
run_message_listenersbefore the listener loop and only visible at debug
logging.- A degraded connection still passes
is_connected()/ ping-pong checks, so
client-side health monitoring cannot detect the server-side pool state.
Proposals (any subset would help)
- Surface
hellometadata (num_connections,approximate_connection_time,
host) anddisconnectreasons via a public callback or at INFO logging. - Emit a loud warning when
num_connectionsinhelloexceeds a threshold
(e.g., 4) while the client manages fewer connections — this is direct
evidence of ghost registrations and imminent interactive-payload loss. - Consider make-before-break reconnects with close-confirmation, or documenting
that reconnect-heavy operation behind NAT can poison the server-side pool. - Still-open PRs #1914 / #1926 address orphaned client-side sessions in the
same failure family — this issue is their server-side counterpart, and since
neither is merged/released (latest release is 3.43.0), apps currently have no
upstream remedy at all; that raises the priority of surfacing the diagnostics
from proposals 1-2.
Happy to share full logs and reproduction notes. We have also filed a parallel
report with Slack developer support regarding the server-side routing/eviction
behavior; will cross-link.
- Ngôn ngữ chính
- Python
- Star
- 4k
- Fork
- 857
- Merge trung bình
- 22 giờ 21 phút
- Pull request đã merge (30 ngày)
- 16
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của slackapi/python-slack-sdk
-
needs info server-side-issue
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
slackapi/python-slack-sdk#1961 · 3 bình luận ·
-
Use logger.isEnabledFor(logging.DEBUG) instead of logger.level <= logging.DEBUG for debug guards Đang mởauto-triage-skip bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
slackapi/python-slack-sdk#1957 ·
-
chat_postMessage silently forwards thread_id to the API, so a threaded reply posts to the channel Đang mởauto-triage-skip enhancement
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
slackapi/python-slack-sdk#1923 · 2 bình luận ·
-
SocketModeClient.connect() retries forever against a permanently closed aiohttp ClientSession Đang mởauto-triage-skip bug socket-mode
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 72/100
slackapi/python-slack-sdk#1922 · 2 bình luận ·
-
auto-triage-skip bug python web-client
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 52/100
slackapi/python-slack-sdk#1853 · 2 bình luận ·
Tất cả issue của slackapi/python-slack-sdk
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
zostera/django-bootstrap4#894 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
use-agent-os/agent-os#3276 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
zephyrproject-rtos/zephyr#119726 ·
-
area/auth bug comp/agent P3 platform/discord type/security
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
NousResearch/hermes-agent#117848 ·