openai / openai/codex

macOS Tahoe: residual native-tls (Secure Transport) connections stall minutes at startup and on quit — coretls TLS-metric ASL logging blocks when diagnosticd is wedged

Open
#38,783 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug CLI connectivity performance
Dominant language
Rust
Stars
125k
Forks
19.5k
PR merge metrics
PR metrics pending

Description

Summary

On macOS Tahoe 26.6, codex (CLI 0.146.0, npm install) intermittently takes 3–5 minutes to run a trivial codex exec "print ok" (~1s of CPU), and interactive sessions hang for minutes on quit after the answer has already been printed. Stack sampling shows the hang is inside Apple's legacy Secure Transport (coretls): every TLS handshake and every TLS context teardown synchronously logs a "TLS metric" through ASL/MessageTracer, and on this machine that logging call blocks on an XPC round-trip to a diagnostics endpoint that is not answering (diagnosticd shows state = not running, never spawned since boot). One blocked thread holds the process-wide ASL mutex, and every other thread that touches TLS queues behind it.

The model call itself succeeds (rustls path), which matches the known mixed TLS backend situation in codex-rs (#6874, #18772): most connections use rustls, but at least one auxiliary connection still goes through reqwest's default native-tls = Secure Transport on macOS. That residual connection is the one that stalls — at startup (handshake metric) and at shutdown (context-destroy metric, producing the quit hang).

Request: finish the rustls migration (default-features = false on reqwest workspace-wide, rustls-tls explicitly everywhere) so no codex connection depends on Secure Transport. That is the same fix already identified in #6874/#18772 for the TLS-1.3 failures; this issue adds a second, more severe reason for it.

Environment

  • macOS Tahoe 26.6 (Darwin 25.6.0), Apple Silicon
  • codex-cli 0.146.0 via npm (@openai/codex, vendored aarch64-apple-darwin binary)
  • Machine condition that triggers it: launchctl print system/com.apple.diagnosticdstate = not running, last exit code = (never exited); analyticsd stuck at spawn scheduled. (This machine-side wedge is an Apple bug — being reported to Apple separately — but codex is the only tool on the machine that hangs because of it, because almost nothing else still uses Secure Transport.)

Measurements (same command, codex exec -s read-only --skip-git-repo-check "print ok" < /dev/null)

condition wall time
healthy window (right after reboot) 6–20 s
degraded (returns within ~1–2 h of any reboot) 28 s → 57 s → 3:00–5:00, monotonically worsening
all hang runs ~1s user CPU; answer often already printed; process alive minutes after

codex mcp list is always instant (no TLS).

Stack evidence (from sample during hangs)

Startup-side stall — a tokio worker inside the handshake, holding the ASL lock, blocked on XPC:

SSLHandshake (Security)
  SSLHandshakeProceed
    tls_handshake_process (libcoretls)
      SSLProcessHandshakeRecordInner
        SSLAdvanceHandshake
          tls_metric_client_finished
            tls_metric_event_add_string
              asl_log (libsystem_asl)
                _asl_lib_vlog
                  _asl_send_message
                    _asl_mt_shim_send_message
                      _xpc_pipe_routine
                        _xpc_pipe_mach_msg
                          mach_msg   <-- blocked here for minutes

Shutdown-side stall (the quit hang) — TLS teardown takes the same path:

_CFRelease (CoreFoundation)
  SSLContextDestroy (Security)
    tls_handshake_destroy (libcoretls)
      tls_metric_event_done
        asl_log
          _asl_evaluate_send
            _pthread_mutex_firstfit_lock_wait   <-- queued behind the thread above

Additional threads (including one in syslog$DARWIN_EXTSN → _vsyslog → _asl_evaluate_send) pile up on the same mutex.

Ruled out by A/B testing (each falsified with interleaved timed runs + mid-hang samples)

  • mcp_servers.node_repl (the ChatGPT-desktop-injected server, 120s startup timeout): disabling it via -c 'mcp_servers.node_repl.enabled=false' — run still stalled 4:55; no node_repl child process is ever spawned during the hang.
  • features.apps=false (the codex_apps streamable-HTTP connector): stalls occur equally in both arms (16.8s / 3:02.9 disabled vs 5.8s / >3:20 control, same sampled stacks).
  • ASL_DISABLE=1: no effect (the MessageTracer shim path is not gated by it) — 4:31.
  • Machine-side remedies short of an OS fix: reboot (relief decays within ~1–2 h), safe-mode boot cycle, unified-log store reset (/var/db/diagnostics, /var/db/uuidtext) — behavior unchanged. No pending macOS update.

Likely related issues

  • #6874, #18772 — prior failures traced to the residual native-tls/Secure Transport backend on macOS; same proposed fix.
  • #14223 (slow exit), #13715 (local clients hang) — plausibly the same mechanism for the macOS reporters, though those issues lack stack samples.

Full sample output files available on request.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by auditing every reqwest dependency declaration across the Rust workspace, using the prior migration context in issues #6874 and #18772. The work is done when workspace-wide reqwest users disable default features and explicitly use rustls, with no codex connection falling back to macOS Secure Transport during startup or shutdown.

Written by the indexing model from the issue text.

Assessment

Tech stack
macos, rust
Domain
cli, networking, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.