Capability by capability: how tests run in parallel and retry, how they're parametrized and set up, how they're described and picked, and how the results come out. Each one with the code that uses it, and its limits where it has them.
01Execution & Resilience
Suites of UI and API tests spend most of their time waiting on the network. Running them in parallel cuts how long a run takes.
-S and -T): 3 suites can run side by side and share a run-wide pool of 12 test threads. Parallelism is declared where it applies:@Suite(parallelized=False) keeps a whole suite to itself.@test(parallelized=False) runs the test alone: nothing else in the run starts until it and its parameters are done.parallelized_parameters=True runs a test's parameters side by side (they run one by one by default).# tj run -S 3 -T 12 (or Runner.run(suite_multithreading_limit=3, # test_multithreading_limit=12)) @Suite() class CheckoutSuite: @test(parameters=CARDS, parallelized_parameters=True) def pays(self, parameter): ... # cards run side by side # runs alone: nothing else starts until it's done @test(parallelized=False) def resets_inventory(self): ...
03Execution & Resilience
Some tests are fine on their own but break each other: one renames the account another one reads, two suites reset the same feature flag. You want everything in parallel except those pairs.
conflicts_with= on a suite or a test and name what it must never overlap with: suites, tests or a mix, as classes or names ("AccountSettingsSuite.renames_account"), even in other files.@Suite(conflicts_with=[AccountSettingsSuite]) # never alongside that suite class LoginSuite: ... @Suite() class ReportsSuite: @test(conflicts_with=[AccountSettingsSuite.renames_account]) def header_shows_account_name(self): ... # never alongside that test
tj audit conflicts lists every pair kept apart without running anything, and the console and JSON report (conflict_waits) show how long a test waited and for which test.04Execution & Resilience
UI and integration tests fail for reasons outside your code. Retrying the right things, and only those, keeps a run green without hiding real failures.
@test(retry=N) retries a test; on a parametrized test only the failing combinations run again.@Suite(retry=N) runs the suite again for its unsuccessful tests and repeats @beforeClass / @afterClass only for the suite parameters (e.g. accounts) that still have failures.tj run --rerun report.json runs only what didn't pass, down to the parameter, and skips setup for accounts with nothing left. It reads an ordinary JSON report: the nightly CI run writes it, and you rerun exactly the failed account × page pairs on your laptop the next morning. From code, Rerun.from_report(path, statuses=[...]) picks which results count.class Flaky(RetryPolicy): when = [...], used as @test(retry=Flaky), @Suite(retry_policy=Flaky) or tj run --retry-policy module:Class: one definition of what to retry, how often and how long to wait, shared by every test that uses it.-p, the console notes the run a test passed on (--flag-flaky labels it FLAKY). Every attempt's traceback is kept, and the HTML report counts retries per test and how much time they cost the run.@Suite(parameters=accounts, retry=2) # suite retry: setup only for class NavigationSuite: # accounts that still failed @test(parameters=pages, retry=3) # per account x page pair def page_opens(self, parameter, suite_parameter): ... # tj run --json-report reports/ then tj run --rerun reports/report.json
--retry N sets every test's retry count for one run; --no-retry turns off test and suite retries, e.g. to see the real failure rate. A test that passed only on a retry is flaky: see Flaky-test detection. A retry policy adds delays, backoff and a separate budget per error (see Exception-type retry control).05Execution & Resilience
A timeout deserves another try. A wrong total on an invoice doesn't. Retrying everything turns real bugs into flaky-looking passes.
retry_on=[...] retries only these exception types, and no_retry_on=[...] never retries these, whatever the retry count. Matching is by class, so subclasses count: retry_on=[requests.exceptions.Timeout] also retries ReadTimeout. When both lists match, no_retry_on wins. The same filters also govern @Suite(retry=) passes, so a suite retry never re-runs a test that failed on a real assertion.For more control, a RetryPolicy:
When(message="503"));delay, backoff and jitter;max_time seconds, or for the rest of the run once circuit=N tests failed despite retrying;@afterClass + @beforeClass again before a retry (reset="class");should_retry() has the last word.@test(retry=3, retry_on=[ConnectionError, TimeoutError], no_retry_on=[AssertionError]) def charges_card(self): ... # network noise: retry; a wrong answer: fail now
from test_junkie.retry import RetryPolicy, When class Flaky(RetryPolicy): when = [When(ConnectionError, attempts=3, delay=1), When(message="503", attempts=2, delay=20, backoff=2)] circuit = 5 # 5 tests still failing: stop @test(retry=Flaky) def refunds_card(self): ...
06Execution & Resilience
A test that failed and then passed on a retry is flaky. Test Junkie says so in the console, the reports and the exit code, so CI can act on it.
retry=N, a retry policy or a suite retry pass.test.is_flaky(parameter, suite_parameter) and get_flaky() name each parameter combination that passed only on a retry, with the run it passed on and what failed before. The JSON report adds "flaky" to every test.tj run --flag-flaky lists them after the summary; --fail-on-flaky also fails the run (exit code 1).<flakyFailure> / <flakyError> on a test that passed in the end, <rerunFailure> / <rerunError> on one that never did, each with the exception's type, message and stack trace. Jenkins reads it with its Flaky Test Handler plugin.<testcase name="charges_card[visa]" ...> <flakyError type="ConnectionError" message="Connection reset by peer: sandbox.payments.example"> <stackTrace>Traceback (most recent call last): ...</stackTrace> </flakyError> </testcase>
from test_junkie.decorators import Suite, test from test_junkie.retry import RetryPolicy, When class Flaky(RetryPolicy): when = [When(ConnectionError, attempts=3, delay=1), When(message="503", attempts=2, delay=20, backoff=2)] circuit = 5 @Suite(feature="Payments", owner="payments-team") class PaymentsApiSuite: @test(parameters=["visa", "amex"], retry=Flaky) def charges_card(self, parameter): # payments sandbox charge = sandbox.charge(card=parameter, amount=1999) assert charge["status"] == "captured"
... Flaky 1 ──────────────────────────────────── PaymentsApiSuite.charges_card [visa] passed on run 2 ConnectionError: Connection reset by peer: sandbox.payments.example FAILED 1 flaky test (--fail-on-flaky) exit code 1
Real output of the suite in this section (the console trimmed to its Flaky block, the XML indented and shortened). With the Flaky Test Handler plugin, Jenkins marks the test flaky, not green. A ConnectionError is an error, so it's <flakyError>; a failed assertion would be <flakyFailure>.
07Test Design
One test, many inputs: every page, every card brand, every API version, each reported on its own.
parameters=pages() calls the function at import and fixes the list then.parameters=pages passes the function itself, and Test Junkie calls it when the run reaches that test, once per run. A test filtered out by tag or name, or skipped, never builds its matrix, so an API call or database query behind it costs nothing on runs that don't need it. The same goes for suite-level parameter functions.str(), so give your objects a readable __str__, or name them with ids=[...] / ids=lambda p: ... (reports and --rerun use those names).def pages(): return [Page(DashboardPage), Page(BillingPage, needs="admin")] @test(parameters=pages) # built when the run reaches the test; def page_opens(self, parameter): ... # filtered out or skipped: never built @test(parameters=pages()) # built at import def page_has_title(self, parameter): ...
1 and "1", or objects with the same __str__) are rejected: the test is ignored with an error naming them. Inside a test, skip("reason") skips just the current parameter, and --rerun repeats only the parameters that failed.08Test Design
Many matrices have two layers: who (account, environment, browser) and what (page, endpoint, card). Setup belongs to the first layer, checks to the second.
@Suite(parameters=accounts) × @test(parameters=pages) runs every page for every account. Each pair is tracked, retried and reported on its own.@beforeClass(suite_parameter), so you log in once per account rather than once per test.retry_on / no_retry_on apply per account × page pair.suite_parameter runs once, not once per account.@Suite(parameters=accounts) class NavigationSuite: @beforeClass() def login(self, suite_parameter): # once per account Browser.login(suite_parameter.email) @test(parameters=pages, parallelized_parameters=True) def page_opens(self, parameter, suite_parameter): parameter.page_object().open() # 2 accounts x 3 pages = 6 runs
09Test Design
Setup and teardown often depend on the suite's parameter: log in as this account, seed data for this tenant. And some tests in a suite shouldn't run the shared per-test setup at all.
@beforeClass, @afterClass, @beforeTest and @afterTest receive the suite parameter just by declaring suite_parameter; hooks that don't need it leave it out.skip_before_test / skip_after_test, and out of the Rules hooks with skip_before_test_rule / skip_after_test_rule: one flag on the test, no checks inside the hook.@afterTest still runs when a test is cancelled mid-way.@beforeTest / @afterTest and Rules can also take a test argument: a read-only view of the test they run around (test.get_tags(), test.get_meta(), test.get_owner()), so one hook can branch per test; nothing a hook does to it reaches the run.@Suite(parameters=["admin", "viewer"]) class AccountSuite: @beforeTest() def open_home(self, suite_parameter): # gets the account Browser.open_home(as_user=suite_parameter) @afterTest() def clear_cookies(self): ... # doesn't need it @test(skip_before_test=True) # starts signed out, skips open_home def signed_out_visit_redirects_to_login(self, suite_parameter): ...
11Metadata & Targeting
When a run of 2,000 tests has 40 failures, the first question is whose they are. Metadata answers it, and drives what runs where.
owner, feature, component, tags and priority are first-class fields on suites and tests.--owners, --features, ...), break down the HTML report and tj audit, and reach listeners.owner= on the suite covers every test, and a test can override it.tj config update), and the run header says when saved config changed a run.@Suite(feature="Checkout", owner="payments-team") class CheckoutSuite: @test(component="cart", tags=["smoke", "ui"], priority=1, meta={"jira": "PAY-123"}) def adds_item(self): ...
meta={...} holds anything else. Not checked: the values themselves, so a misspelled tag goes unnoticed.12Metadata & Targeting
Tests without an owner rot. An audit shows who owns what and where the gaps are, and a CI gate keeps new tests from arriving without one.
tj audit breaks your tests down by suite, owner, feature, component or tag without running them, with each group's share of the tests audited.--fail-on-gaps fails CI when tests are missing an owner, feature, component or tag (or just the ones you list).--json feeds dashboards.--no-owners, --no-tags or --no-test-retries, and the gaps section prints the exact command that lists the offenders.# who owns what, without running anything tj audit owners -s tests # in CI: fail when a test has no owner or feature tj audit owners -s tests --fail-on-gaps owners,features # which Checkout tests have no tag, by component tj audit components -s tests --features Checkout --no-tags
13Metadata & Targeting
Run the smoke checks and the slowest suites first: failures surface sooner, and the long tail doesn't hold up the end of the run.
priority= on suites and tests; lower runs first.@Suite(priority=1) class SmokeSuite: ... @Suite(priority=2) class CheckoutSuite: @test(priority=1) def adds_item(self): ...
14Metadata & Targeting
Shuffling tests flushes out hidden dependencies between them, but shuffling everything at once makes every run different.
@Suite(order=...) picks ALPHABETICAL, RANDOM, PRIORITY_ASC or PRIORITY_DESC for one suite.--seed N repeats a random order.from test_junkie.constants import TestOrder @Suite(order=TestOrder.RANDOM) # shuffle just this suite class CartSuite: ... # tj run --seed 4242 repeat yesterday's order
15Metadata & Targeting
CI runs smoke tests on every push and the full set at night; an owner wants just their tests. Picking tests should take a flag, not a script.
--tags-any / --tags-all and their --skip-tags-* pair.--owners, --components and --features, each a flag of its own.--tests (names, Suite.test or patterns) and -x for suites.Runner.run() (tags through tag_config), and can be saved in the project config.tj run --tags-any smoke --skip-tags-any slow
tj run --owners payments-team --features Checkout
tj run -t "CheckoutSuite.pays_*"
16Metadata & Targeting
Some facts are only known while the test runs: which order it created, which known bug it hit.
Meta.update(order_id=order.id) in a running test, any helper it calls, or @beforeTest / @afterTest sets metadata for exactly the parameter, suite parameter and attempt that's running, with no self to pass.Meta.get() reads it back and Meta.bind() carries it into threads the test starts.Meta.link(label, url) and Meta.attach(name, bytes_or_path) add links and files.properties["test_meta"] (per attempt in properties["meta_attempts"]).@test(parameters=CARDS) def charges_card(self, parameter): order = shop.checkout(parameter) Meta.update(order_id=order.id)
meta per test and meta_set per run; the HTML report shows a metadata panel with clickable links and attachments embedded up to 512 KB; the XML report writes JUnit <properties>. Owner, tags and component can't change at run time.17Results & Reporting
When login breaks in setup, a report that says "41 failed" sends 41 people looking at the wrong thing. The useful answer is "1 setup error, 40 tests never ran".
AssertionError is a failure (the product is wrong); any other exception is an error (the test or the environment broke). The same split applies to @beforeClass, @afterClass and group hooks, each with its own listener event.@beforeClass or @beforeGroup failed. They're recorded per parameter, each carrying the error that stopped it, so with two accounts only the account whose login broke is ignored.Runner.cancel() stopped before they started. Reports are still written.@Suite(parameters=["admin", "viewer"]) class BillingSuite: @beforeClass() def login(self, suite_parameter): ... # raises for "viewer" @test(parameters=pages) def page_opens(self, parameter, suite_parameter): ... # admin x pages: run normally # viewer x pages: ignored, each one carrying the login error
--rerun pick them up.18Results & Reporting
People read HTML, CI reads XML, scripts read JSON.
--html-report, --xml-report, --json-report (a folder works too). With -m, CPU and memory are charted in the HTML report and in the console. Screenshots and other files from Meta.attach() are embedded in the HTML report (up to 512 KB each, bigger ones saved next to it). See a sample HTML report.tj run -m --html-report reports/ --xml-report reports/ --json-report reports/
--rerun reads.classname and time per test, parameters in the test case name (totals[viewer]), the exception's type, message and <stackTrace>, retried runs recorded the Maven Surefire way (<flakyFailure> / <rerunFailure>), and metadata as <properties>.19Results & Reporting
Results belong in your own systems as they happen: a database, Slack, a dashboard.
Listener class per suite (@Suite(listener=...)), with 22 named events: 9 per test (including on_retry), 9 per suite and 4 for group hooks, including separate events for a @beforeClass failure versus error.on_failure fires on every attempt and on_complete once with the final result, so one listener can log flakiness and another record the outcome.class ResultsToDb(Listener): def on_complete(self, **kwargs): # once per test x parameter test = kwargs["properties"]["jm"]["jto"] account = kwargs["properties"]["suite_meta"]["parameter"] db.insert(test=test.get_function_name(), account=account)
run() raises once reports are written. There's no run-wide listener or run start/end event; share one class across suites.20Results & Reporting
After a run, scripts decide what happens next: open tickets, compare with yesterday, gate a deploy.
Runner.run() returns an aggregator with the run's statistics, already grouped: get_report_by_owner(), get_report_by_features() and get_report_by_tags(), so routing failures to each team is one call. If the run itself errored, it raises after writing reports instead.runner.get_executed_suites() always returns the SuiteObjects and TestObjects, down to each parameter, suite parameter and attempt: status, timing, retries and the policy behind them, metadata per attempt, hook timings, captured output, owners and tags.isinstance works.runner = Runner([CheckoutSuite]) runner.run() for suite in runner.get_executed_suites(): for test in suite.get_test_objects(): print(test.get_function_name(), test.get_owner())
Rerun().add(...) and pass it straight back to Runner.run(rerun=...). test.is_flaky() / get_flaky() say which parameters passed only on a retry.