Retrying flaky tests on specific exceptions
A blanket retry hides real bugs behind flaky infrastructure. retry_on and no_retry_on scope retries to exception types: a connection timeout runs again, a wrong response fails on the first run, the way it should.
How retry= counts
retry=N is the total number of runs: retry=2 runs a failing test at most twice (one retry). A test that passes is never run again.
The problem with plain retry=
On its own, retry= retries after any exception, assertion failures included. A real contract break gets run again, and either passes by luck or just fails twice:
@test(retry=2) # retries after anything - including real bugs def response_shape(self): data = api.get("/users/1") # if the API renamed "name", this AssertionError runs again before the test fails assert "name" in data
The allowlist: retry_on=
retry_on=[...] retries only when the exception is one of those types, or a subclass of one. Anything else fails at once:
import requests from test_junkie.decorators import Suite, test @Suite() class UserApiSuite: @test( retry=3, retry_on=[ requests.exceptions.ConnectionError, requests.exceptions.Timeout, # also ReadTimeout and ConnectTimeout, its subclasses ], ) def response_shape(self): data = api.get("/users/1") # AssertionError isn't in retry_on: a wrong shape fails once, honestly assert "name" in data
retry_on=[requests.exceptions.Timeout] also retries a ReadTimeout, which is what requests actually raises. Before 0.9a6 only the exact listed type matched, so list the subclasses too if you're on an older version.The blocklist: no_retry_on=
Sometimes the exceptions not to retry are easier to name. no_retry_on retries everything except those types (and their subclasses):
@test(retry=2, no_retry_on=[AssertionError, ValueError]) def status_endpoint(self): # runs again after connection errors, timeouts, DNS failures... # never after an assertion or a ValueError response = requests.get("https://api.example.com/status", timeout=5) assert response.status_code == 200
A test can use both lists: no_retry_on is checked first, then retry_on. Usually one is enough: a short allowlist when few exceptions are worth a retry, a short blocklist when few aren't.
A real API, end to end
The same idea against a public API you can run right now, httpbin.org:
import requests from test_junkie.decorators import Suite, test TRANSIENT = [requests.exceptions.ConnectionError, requests.exceptions.Timeout] @Suite() class HttpBinContractSuite: @test(retry=3, retry_on=TRANSIENT) def status_endpoint_is_reachable(self): response = requests.get("https://httpbin.org/status/200", timeout=5) assert response.status_code == 200 @test(retry=3, retry_on=TRANSIENT) def response_has_expected_shape(self): data = requests.get("https://httpbin.org/get", timeout=5).json() assert "url" in data # a broken contract fails on the first run
tj run -s test_httpbin_contract.py
Retries with parameters
When a test has parameters= and retry=, retries are per variant. If the safari variant times out, only it runs again:
@test( retry=2, retry_on=[ConnectionError, TimeoutError], parameters=[{"browser": "chrome"}, {"browser": "firefox"}, {"browser": "safari"}], ) def checkout_works(self, parameter): ...
Overriding retries for one run
From the command line, tj run --no-retry runs every test once (to see the real failure rate locally) and tj run --retry 3 gives every failing test up to 3 runs (e.g. in CI against a shaky environment), whatever its retry= says. retry_on and no_retry_on still apply.