Tutorials / Core features

Retrying flaky tests on specific exceptions

A blanket retry hides real bugs behind flaky infrastructure. retry_on and no_retry_on scope retries to exception types: a connection timeout runs again, a wrong response fails on the first run, the way it should.

How retry= counts

retry=N is the total number of runs: retry=2 runs a failing test at most twice (one retry). A test that passes is never run again.

The problem with plain retry=

On its own, retry= retries after any exception, assertion failures included. A real contract break gets run again, and either passes by luck or just fails twice:

test_api.pypython
@test(retry=2)  # retries after anything - including real bugs
def response_shape(self):
    data = api.get("/users/1")
    # if the API renamed "name", this AssertionError runs again before the test fails
    assert "name" in data

The allowlist: retry_on=

retry_on=[...] retries only when the exception is one of those types, or a subclass of one. Anything else fails at once:

test_api.pypython
import requests
from test_junkie.decorators import Suite, test


@Suite()
class UserApiSuite:

    @test(
        retry=3,
        retry_on=[
            requests.exceptions.ConnectionError,
            requests.exceptions.Timeout,  # also ReadTimeout and ConnectTimeout, its subclasses
        ],
    )
    def response_shape(self):
        data = api.get("/users/1")
        # AssertionError isn't in retry_on: a wrong shape fails once, honestly
        assert "name" in data
Subclasses count (since 0.9a6)
retry_on=[requests.exceptions.Timeout] also retries a ReadTimeout, which is what requests actually raises. Before 0.9a6 only the exact listed type matched, so list the subclasses too if you're on an older version.

The blocklist: no_retry_on=

Sometimes the exceptions not to retry are easier to name. no_retry_on retries everything except those types (and their subclasses):

test_api.pypython
@test(retry=2, no_retry_on=[AssertionError, ValueError])
def status_endpoint(self):
    # runs again after connection errors, timeouts, DNS failures...
    # never after an assertion or a ValueError
    response = requests.get("https://api.example.com/status", timeout=5)
    assert response.status_code == 200

A test can use both lists: no_retry_on is checked first, then retry_on. Usually one is enough: a short allowlist when few exceptions are worth a retry, a short blocklist when few aren't.

A real API, end to end

The same idea against a public API you can run right now, httpbin.org:

test_httpbin_contract.pypython
import requests
from test_junkie.decorators import Suite, test

TRANSIENT = [requests.exceptions.ConnectionError, requests.exceptions.Timeout]


@Suite()
class HttpBinContractSuite:

    @test(retry=3, retry_on=TRANSIENT)
    def status_endpoint_is_reachable(self):
        response = requests.get("https://httpbin.org/status/200", timeout=5)
        assert response.status_code == 200

    @test(retry=3, retry_on=TRANSIENT)
    def response_has_expected_shape(self):
        data = requests.get("https://httpbin.org/get", timeout=5).json()
        assert "url" in data  # a broken contract fails on the first run
terminalbash
tj run -s test_httpbin_contract.py

Retries with parameters

When a test has parameters= and retry=, retries are per variant. If the safari variant times out, only it runs again:

checkout_suite.pypython
@test(
    retry=2,
    retry_on=[ConnectionError, TimeoutError],
    parameters=[{"browser": "chrome"}, {"browser": "firefox"}, {"browser": "safari"}],
)
def checkout_works(self, parameter):
    ...

Overriding retries for one run

From the command line, tj run --no-retry runs every test once (to see the real failure rate locally) and tj run --retry 3 gives every failing test up to 3 runs (e.g. in CI against a shaky environment), whatever its retry= says. retry_on and no_retry_on still apply.