A nightly API job usually fails in one of two ways: loudly enough that the scheduler marks it red, or quietly enough that nobody notices until tomorrow’s report is missing. The quiet version often comes from one innocent line of Python: requests.get(url) with no timeout and no bounded retry policy.
That line is fine for a shell experiment. It is weak plumbing for a job that pulls paginated filings, Treasury datasets, webhook receipts, or any other source where the network can stall. The fix is not a pile of sleeps around the call. It is a small Requests session with a connect timeout, a read timeout, status-aware retries, and a rule for which HTTP methods are safe to repeat.
On this page
Give every request two timeouts, not one vague number#
The Requests timeout documentation is blunt about the default: if you do not pass a timeout, Requests does not time out. The same section also notes that timeout is not a total wall-clock limit for the whole response body; it is about how long the client waits for socket activity.
For API clients, a tuple is clearer than a single float. The Requests API reference defines timeout as either a float or a (connect timeout, read timeout) tuple. The connect timeout covers the attempt to open the connection. The read timeout covers waiting for the server to send data after the request is in flight.
import requests
session = requests.Session()
response = session.get(
"https://api.example.com/v1/orders",
timeout=(3.05, 20),
)
response.raise_for_status()
payload = response.json()
That tuple says: give the network a few seconds to connect, then give the API a longer but still finite window to produce bytes. Pick values from the job’s actual tolerance. A dashboard refresh may want a short read timeout. A large export endpoint may deserve more time, but it still needs an upper bound.
Retry connection failures differently from read failures#
Retries are not one bucket. A connect failure happens before the request reaches the remote service, so it is usually safe to repeat. Requests documents ConnectTimeout as safe to retry in its exception reference on the API page. A read timeout is different: the server may already have received the request, started work, or changed state before the client gave up waiting.
That distinction matters when a script calls a write endpoint. Retrying a GET for a public dataset is usually fine. Retrying a POST that creates an invoice, sends an email, or places an order is not safe unless the API gives you an idempotency key and documents the behavior.
The clean pattern is to put default retries on idempotent reads and leave state-changing requests out unless the API has a specific idempotency contract.
Mount an HTTPAdapter with a Retry policy#
Requests uses adapters for transport settings. The Requests API reference documents HTTPAdapter, and retry behavior is configured with urllib3’s Retry object. The urllib3 Retry reference says the default allowed methods are uppercase idempotent methods such as GET, HEAD, PUT, DELETE, OPTIONS, and TRACE. It also documents status_forcelist, backoff_factor, backoff_jitter, and respect_retry_after_header.
from requests import Session
from requests.adapters import HTTPAdapter
from urllib3.util import Retry
retry = Retry(
total=4,
connect=4,
read=0,
status=4,
allowed_methods={"GET", "HEAD", "OPTIONS"},
status_forcelist={429, 500, 502, 503, 504},
backoff_factor=0.5,
backoff_jitter=0.25,
respect_retry_after_header=True,
)
adapter = HTTPAdapter(max_retries=retry)
session = Session()
session.mount("https://", adapter)
session.mount("http://", adapter)
This policy makes four choices explicit. It retries connection failures. It does not retry read errors, because those happen after the request may have been sent. It retries selected transient status codes only for selected methods. It also uses backoff and respects Retry-After when the server sends one for supported retry-after statuses.
Do not let JSON parsing hide a failed response#
A common API bug is to call response.json() and treat a parsed object as proof of success. The Requests quickstart warns against that: JSON decoding can succeed even when the HTTP status represents an error. Some APIs return a neat JSON error envelope with a 429 or 500 status.
Call raise_for_status() before trusting the payload. If the job needs structured error logging, catch requests.HTTPError around that call and record the status code plus a short, privacy-clean body preview. Do not log credentials, bearer tokens, private hostnames, or full request headers.
import requests
try:
response = session.get(
"https://api.example.com/v1/reports/latest",
timeout=(3.05, 20),
)
response.raise_for_status()
data = response.json()
except requests.Timeout as exc:
raise RuntimeError("API request timed out") from exc
except requests.HTTPError as exc:
status = exc.response.status_code if exc.response else "unknown"
raise RuntimeError(f"API request failed with HTTP {status}") from exc
This does not make the network perfect. It makes failure visible and bounded, which is what a scheduled job needs.
Use a session per API, not global retry magic#
Different APIs deserve different patience. A small status endpoint should not share the same read timeout as a bulk filing download. A rate-limited government dataset should respect server pacing; a private webhook verifier may need to fail fast so the caller can retry.
For finance and public-data jobs, this pattern pairs well with source-specific clients. If you are pulling SEC data, keep the timeout and retry rules next to the code described in the SEC EDGAR XBRL API guide. If you are polling Treasury data, do the same around the calls in the Treasury FiscalData API guide. The API client becomes the one place where pacing, headers, status handling, and JSON parsing rules live.
A practical default for read-only API jobs#
For a read-only scheduled job, start with this rule set: timeout=(3.05, 20), a retry budget of three or four attempts, retries limited to GET and HEAD, a status list that includes 429, 500, 502, 503, and 504, plus backoff with a little jitter. Then tune from logs rather than guesses.
Keep write calls separate. If an API documents idempotency keys, use them and test the failure mode in a sandbox before enabling retries. If it does not, make the write fail once, record enough context to reconcile it, and let the scheduler alert a human instead of guessing whether a second attempt is safe.
The next time you touch a Python API job, search for bare requests.get(, requests.post(, and requests.request(. Add explicit timeouts first. Then add a mounted retry policy only where the HTTP method and API contract make repeat attempts safe.
Leave a Reply