Retry: sleep_for_retry uses max(wait, delay_max) — inflates small Retry-After to delay_max (60s floor)

Open Beginner friendly
#860 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
76/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Quiet
Tech stack
python
Domain
backend

Research direction

Read src/databricks/sql/auth/retry.py, starting with DatabricksRetryPolicy.sleep_for_retry and the neighboring get_backoff_time clamp. Reproduce the small Retry-After case and inspect tests/test_error_recovery.py, especially the progressive Retry-After scenario. Done means Retry-After values are capped rather than raised to delay_max, with the affected retry paths no longer waiting for the 60-second floor.

Written by the indexing model from the issue text.

Description

engineer-bot

Summary

DatabricksRetryPolicy.sleep_for_retry applies delay_max as a floor instead of a ceiling, so a small server Retry-After (e.g. 2) is inflated to the full delay_max (default 60s) on every retry. This makes retryable errors that advertise a short Retry-After sleep far longer than the server requested.

Location

src/databricks/sql/auth/retry.py, in sleep_for_retry:

retry_after = self.get_retry_after(response)
if retry_after:
    proposed_wait = retry_after
else:
    proposed_wait = self.get_backoff_time()

proposed_wait = max(proposed_wait, self.delay_max)   # <-- BUG: floor, not ceiling
...
time.sleep(proposed_wait)

max(proposed_wait, self.delay_max) guarantees the sleep is at least delay_max. With the default _retry_delay_max = 60, a server response of Retry-After: 2 results in max(2, 60) = 60s.

Why it's a bug

The sibling method get_backoff_time in the same file does the opposite (and correct) clamp, with a docstring that states the intent:

# get_backoff_time():
#   "Never returns a value larger than self.delay_max"
proposed_backoff = min(proposed_backoff, self.delay_max)

So delay_max is intended as a ceiling on the wait. sleep_for_retry inverts it. The fix is to cap (not floor) the proposed wait — min(proposed_wait, self.delay_max) — or to not clamp an explicit server Retry-After upward at all.

Impact / repro

A server that returns 503 with a small Retry-After (say 2s, then 4s, then success) is honored as 60s, then 60s — 120s total instead of the intended ~6s.

Observed in the driver-test conformance suite (ERRORRECOV-001, "HTTP 503 with progressive Retry-After"): the request-executing test blocked in retry.py sleep_for_retry -> time.sleep(60) twice and hit the 120s pytest-timeout. faulthandler stack (SEA backend):

tests/test_error_recovery.py:70 cur.execute(SIMPLE_QUERY)
  -> databricks/sql/backend/sea/backend.py execute_command
  -> .../sea/utils/http_client.py _make_request
  -> urllib3 connectionpool.urlopen -> retries.sleep(response)
  -> databricks/sql/auth/retry.py:~301 sleep_for_retry -> time.sleep(proposed_wait)

Backend-agnostic: it's in the shared DatabricksRetryPolicy, so both the SEA and Thrift HTTP paths are affected (the kernel path uses a different retry mechanism and is unaffected).

Suggested fix

proposed_wait = min(proposed_wait, self.delay_max)

(and confirm delay_max is the intended upper bound on an honored Retry-After, matching get_backoff_time).

Dominant language
Python
Stars
233
Forks
152
Avg merge
21h 5m
Merged PRs (30d)
10

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from databricks/databricks-sql-python

All issues in databricks/databricks-sql-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.