Back to home
Features Test Junkie 0.9a8

What Test Junkie does

Capability by capability: how tests run in parallel and retry, how they're parametrized and set up, how they're described and picked, and how the results come out. Each one with the code that uses it, and its limits where it has them.

01Execution & Resilience

Parallel execution

Suites of UI and API tests spend most of their time waiting on the network. Running them in parallel cuts how long a run takes.

Test Junkie runs tests on threads. The number of suites and the number of tests that run at once are two separate limits (-S and -T): 3 suites can run side by side and share a run-wide pool of 12 test threads. Parallelism is declared where it applies:
  • @Suite(parallelized=False) keeps a whole suite to itself.
  • @test(parallelized=False) runs the test alone: nothing else in the run starts until it and its parameters are done.
  • parallelized_parameters=True runs a test's parameters side by side (they run one by one by default).
Threads share one process, so one logged-in session or one database connection pool can be shared across tests. The trade-off: CPU-heavy tests don't speed up, because threads share one core under the GIL.
example.pypython
# tj run -S 3 -T 12        (or Runner.run(suite_multithreading_limit=3,
#                                         test_multithreading_limit=12))
@Suite()
class CheckoutSuite:

    @test(parameters=CARDS, parallelized_parameters=True)
    def pays(self, parameter): ...          # cards run side by side

    # runs alone: nothing else starts until it's done
    @test(parallelized=False)
    def resets_inventory(self): ...

↑ Back to the list

02Execution & Resilience

Shared-resource control

Parallel runs are fast until they're too fast for something tests share: a Selenium Grid with five browser slots, a sandbox API that allows two calls a second, a staging database that falls over at twenty connections. Test Junkie caps the tests that use the resource and leaves the rest of the run at full speed.

Three tools, from targeted to broad:
  • Resource pools cap only the tests that use a resource. Limiter.pool("grid", max_concurrent=5) plus uses="grid" on a test or suite: at most 5 browser tests at once, while API tests keep running. A slot covers the test with its @beforeTest / @afterTest and is released between retry attempts. min_interval spaces out starts for rate-limited APIs. Tests using several pools take them in a fixed order, so they can't deadlock. A misspelled pool name (uses="gird") stops the run before anything starts. Ctrl+C while a test waits for a slot cancels it cleanly instead of hanging.
  • Throttling sets the minimum time between two starts across the whole run, with @Suite(throttling=) to exempt or slow one suite.
  • Ramp-up grows the thread limits from 1 to -S/-T over N seconds, for grids that autoscale or services that need warming.
example.pypython
from test_junkie.objects import Limiter

Limiter.pool("grid", max_concurrent=5)                   # 5 browsers
Limiter.pool("payments", max_concurrent=2, min_interval=0.5)

@Suite(uses="grid")                  # every test takes a grid slot
class CheckoutUiSuite: ...

@Suite(throttling=0)                 # exempt from run-wide throttling
class PricingApiSuite: ...

# tj run -T 20 --test-throttling 0.5 --ramp-up 30
Limits
Pools are defined in code; throttling and ramp-up also work from tj run flags or a saved config. A test waiting for a pool slot holds one of the -T threads while it waits, and pools coordinate one process, not several machines.

↑ Back to the list

03Execution & Resilience

Parallel conflict rules

Some tests are fine on their own but break each other: one renames the account another one reads, two suites reset the same feature flag. You want everything in parallel except those pairs.

Put conflicts_with= on a suite or a test and name what it must never overlap with: suites, tests or a mix, as classes or names ("AccountSettingsSuite.renames_account"), even in other files.
  • That one line is the whole rule. It applies in both directions, so the other side needs no change.
  • Everything not named keeps running in parallel, and a suite held back by a conflict doesn't block the queue: the scheduler starts the next suites and comes back to it.
  • Every target is checked before the run starts, and a wrong one stops the run with a link to the docs.
example.pypython
@Suite(conflicts_with=[AccountSettingsSuite])   # never alongside that suite
class LoginSuite: ...

@Suite()
class ReportsSuite:

    @test(conflicts_with=[AccountSettingsSuite.renames_account])
    def header_shows_account_name(self): ...   # never alongside that test
See what was kept apart
tj audit conflicts lists every pair kept apart without running anything, and the console and JSON report (conflict_waits) show how long a test waited and for which test.

↑ Back to the list

04Execution & Resilience

Test retries

UI and integration tests fail for reasons outside your code. Retrying the right things, and only those, keeps a run green without hiding real failures.

  • @test(retry=N) retries a test; on a parametrized test only the failing combinations run again.
  • @Suite(retry=N) runs the suite again for its unsuccessful tests and repeats @beforeClass / @afterClass only for the suite parameters (e.g. accounts) that still have failures.
  • Rerun from a file, anywhere. tj run --rerun report.json runs only what didn't pass, down to the parameter, and skips setup for accounts with nothing left. It reads an ordinary JSON report: the nightly CI run writes it, and you rerun exactly the failed account × page pairs on your laptop the next morning. From code, Rerun.from_report(path, statuses=[...]) picks which results count.
  • Retry policies. class Flaky(RetryPolicy): when = [...], used as @test(retry=Flaky), @Suite(retry_policy=Flaky) or tj run --retry-policy module:Class: one definition of what to retry, how often and how long to wait, shared by every test that uses it.
  • Retries stay visible. With -p, the console notes the run a test passed on (--flag-flaky labels it FLAKY). Every attempt's traceback is kept, and the HTML report counts retries per test and how much time they cost the run.
example.pypython
@Suite(parameters=accounts, retry=2)       # suite retry: setup only for
class NavigationSuite:                     # accounts that still failed

    @test(parameters=pages, retry=3)       # per account x page pair
    def page_opens(self, parameter, suite_parameter): ...

# tj run --json-report reports/   then   tj run --rerun reports/report.json
From the command line
--retry N sets every test's retry count for one run; --no-retry turns off test and suite retries, e.g. to see the real failure rate. A test that passed only on a retry is flaky: see Flaky-test detection. A retry policy adds delays, backoff and a separate budget per error (see Exception-type retry control).

↑ Back to the list

05Execution & Resilience

Exception-type retry control

A timeout deserves another try. A wrong total on an invoice doesn't. Retrying everything turns real bugs into flaky-looking passes.

retry_on=[...] retries only these exception types, and no_retry_on=[...] never retries these, whatever the retry count. Matching is by class, so subclasses count: retry_on=[requests.exceptions.Timeout] also retries ReadTimeout. When both lists match, no_retry_on wins. The same filters also govern @Suite(retry=) passes, so a suite retry never re-runs a test that failed on a real assertion.

For more control, a RetryPolicy:

  • also matches the message (When(message="503"));
  • gives each condition its own budget, delay, backoff and jitter;
  • can stop retrying after max_time seconds, or for the rest of the run once circuit=N tests failed despite retrying;
  • can run @afterClass + @beforeClass again before a retry (reset="class");
  • and should_retry() has the last word.
example.pypython
@test(retry=3,
      retry_on=[ConnectionError, TimeoutError],
      no_retry_on=[AssertionError])
def charges_card(self): ...
# network noise: retry; a wrong answer: fail now
retry_policy.pypython
from test_junkie.retry import RetryPolicy, When

class Flaky(RetryPolicy):
    when = [When(ConnectionError, attempts=3, delay=1),
            When(message="503", attempts=2,
                 delay=20, backoff=2)]
    circuit = 5        # 5 tests still failing: stop

@test(retry=Flaky)
def refunds_card(self): ...

↑ Back to the list

06Execution & Resilience

Flaky-test detection & CI reporting

A test that failed and then passed on a retry is flaky. Test Junkie says so in the console, the reports and the exit code, so CI can act on it.

Flaky is tracked whatever retried the test: retry=N, a retry policy or a suite retry pass.
  • test.is_flaky(parameter, suite_parameter) and get_flaky() name each parameter combination that passed only on a retry, with the run it passed on and what failed before. The JSON report adds "flaky" to every test.
  • tj run --flag-flaky lists them after the summary; --fail-on-flaky also fails the run (exit code 1).
  • The XML report records every failed run the Maven Surefire way: <flakyFailure> / <flakyError> on a test that passed in the end, <rerunFailure> / <rerunError> on one that never did, each with the exception's type, message and stack trace. Jenkins reads it with its Flaky Test Handler plugin.
reports/report.xmlxml
<testcase name="charges_card[visa]" ...>
  <flakyError type="ConnectionError"
      message="Connection reset by peer: sandbox.payments.example">
    <stackTrace>Traceback (most recent call last): ...</stackTrace>
  </flakyError>
</testcase>
tests/payments_api.pypython
from test_junkie.decorators import Suite, test
from test_junkie.retry import RetryPolicy, When

class Flaky(RetryPolicy):
    when = [When(ConnectionError, attempts=3, delay=1),
            When(message="503", attempts=2, delay=20, backoff=2)]
    circuit = 5

@Suite(feature="Payments", owner="payments-team")
class PaymentsApiSuite:

    @test(parameters=["visa", "amex"], retry=Flaky)
    def charges_card(self, parameter):       # payments sandbox
        charge = sandbox.charge(card=parameter, amount=1999)
        assert charge["status"] == "captured"
tj run -s tests --fail-on-flaky --xml-report reports/output
...
Flaky 1 ────────────────────────────────────

  PaymentsApiSuite.charges_card [visa]  passed on run 2
      ConnectionError: Connection reset by peer: sandbox.payments.example

 FAILED   1 flaky test (--fail-on-flaky)  exit code 1

Real output of the suite in this section (the console trimmed to its Flaky block, the XML indented and shortened). With the Flaky Test Handler plugin, Jenkins marks the test flaky, not green. A ConnectionError is an error, so it's <flakyError>; a failed assertion would be <flakyFailure>.

↑ Back to the list

07Test Design

Parametrized execution

One test, many inputs: every page, every card brand, every API version, each reported on its own.

You choose when a test's matrix is built:
  • parameters=pages() calls the function at import and fixes the list then.
  • parameters=pages passes the function itself, and Test Junkie calls it when the run reaches that test, once per run. A test filtered out by tag or name, or skipped, never builds its matrix, so an API call or database query behind it costs nothing on runs that don't need it. The same goes for suite-level parameter functions.
Any Python object works as a parameter: dicts, dataclasses, page-object classes. Each one is reported by its str(), so give your objects a readable __str__, or name them with ids=[...] / ids=lambda p: ... (reports and --rerun use those names).
example.pypython
def pages():
    return [Page(DashboardPage), Page(BillingPage, needs="admin")]

@test(parameters=pages)        # built when the run reaches the test;
def page_opens(self, parameter): ...   # filtered out or skipped: never built

@test(parameters=pages())      # built at import
def page_has_title(self, parameter): ...
Good to know
Two parameters that read the same (1 and "1", or objects with the same __str__) are rejected: the test is ignored with an error naming them. Inside a test, skip("reason") skips just the current parameter, and --rerun repeats only the parameters that failed.

↑ Back to the list

08Test Design

Suite-level × test-level cross-multiply

Many matrices have two layers: who (account, environment, browser) and what (page, endpoint, card). Setup belongs to the first layer, checks to the second.

@Suite(parameters=accounts) × @test(parameters=pages) runs every page for every account. Each pair is tracked, retried and reported on its own.
  • The account reaches @beforeClass(suite_parameter), so you log in once per account rather than once per test.
  • A suite retry repeats setup only for the accounts that still failed, and retry_on / no_retry_on apply per account × page pair.
  • A test that doesn't take suite_parameter runs once, not once per account.
Limits: a test's parameter function is built once per run, not per account, and hooks never see the test-level parameter.
example.pypython
@Suite(parameters=accounts)
class NavigationSuite:

    @beforeClass()
    def login(self, suite_parameter):        # once per account
        Browser.login(suite_parameter.email)

    @test(parameters=pages, parallelized_parameters=True)
    def page_opens(self, parameter, suite_parameter):
        parameter.page_object().open()       # 2 accounts x 3 pages = 6 runs

↑ Back to the list

09Test Design

Parameter-aware lifecycle hooks

Setup and teardown often depend on the suite's parameter: log in as this account, seed data for this tenant. And some tests in a suite shouldn't run the shared per-test setup at all.

  • @beforeClass, @afterClass, @beforeTest and @afterTest receive the suite parameter just by declaring suite_parameter; hooks that don't need it leave it out.
  • Within the same suite, any test can opt out of the per-test hooks with skip_before_test / skip_after_test, and out of the Rules hooks with skip_before_test_rule / skip_after_test_rule: one flag on the test, no checks inside the hook.
  • @afterTest still runs when a test is cancelled mid-way.
  • @beforeTest / @afterTest and Rules can also take a test argument: a read-only view of the test they run around (test.get_tags(), test.get_meta(), test.get_owner()), so one hook can branch per test; nothing a hook does to it reaches the run.
example.pypython
@Suite(parameters=["admin", "viewer"])
class AccountSuite:

    @beforeTest()
    def open_home(self, suite_parameter):     # gets the account
        Browser.open_home(as_user=suite_parameter)

    @afterTest()
    def clear_cookies(self): ...              # doesn't need it

    @test(skip_before_test=True)       # starts signed out, skips open_home
    def signed_out_visit_redirects_to_login(self, suite_parameter): ...
Limits
Hooks receive the suite parameter only, never a test's own parameter, and hooks don't declare dependencies on one another.

↑ Back to the list

10Test Design

Shared lifecycle across suites

Some setup spans several suites: seed shared accounts once for three suites, start a mock server for the API suites only, clean up after the last of them finishes.

@GroupRules names any set of suites, from any files, and attaches @beforeGroup (before the first of them starts) and @afterGroup (after the last finishes), so cleanup happens as soon as that set of suites is done. The group is made of the suites passed to the Runner.
  • If @beforeGroup fails, the group's tests are recorded as ignored instead of failing one by one, and group hooks fire their own listener events.
  • With -S above 1, @beforeGroup runs once and the group's other suites wait for it.
  • @afterGroup runs after the last suite of the group however it ended (passed, skipped, filtered out, ignored or cancelled) and after a failed @beforeGroup; it doesn't run if @beforeGroup never ran.
example.pypython
@GroupRules()
def shared_accounts(rules):

    @beforeGroup([AccountSettingsSuite, LoginSuite, ReportsSuite])
    def seed():
        db.load("accounts.sql")

    @afterGroup([AccountSettingsSuite, LoginSuite, ReportsSuite])
    def clean_up():
        db.truncate("accounts")
Per suite, too
Rules classes add before/after hooks to every suite that uses them.

↑ Back to the list

11Metadata & Targeting

Structured test metadata

When a run of 2,000 tests has 40 failures, the first question is whose they are. Metadata answers it, and drives what runs where.

owner, feature, component, tags and priority are first-class fields on suites and tests.
  • A wrong type fails when the class loads, naming the argument.
  • They filter runs (--owners, --features, ...), break down the HTML report and tj audit, and reach listeners.
  • One owner= on the suite covers every test, and a test can override it.
  • A team can save its filters in the project config (tj config update), and the run header says when saved config changed a run.
example.pypython
@Suite(feature="Checkout", owner="payments-team")
class CheckoutSuite:

    @test(component="cart", tags=["smoke", "ui"], priority=1,
          meta={"jira": "PAY-123"})
    def adds_item(self): ...
Good to know
meta={...} holds anything else. Not checked: the values themselves, so a misspelled tag goes unnoticed.

↑ Back to the list

12Metadata & Targeting

Test inventory audit & CI gate

Tests without an owner rot. An audit shows who owns what and where the gaps are, and a CI gate keeps new tests from arriving without one.

tj audit breaks your tests down by suite, owner, feature, component or tag without running them, with each group's share of the tests audited.
  • --fail-on-gaps fails CI when tests are missing an owner, feature, component or tag (or just the ones you list).
  • --json feeds dashboards.
  • It combines with the run filters and with gap filters such as --no-owners, --no-tags or --no-test-retries, and the gaps section prints the exact command that lists the offenders.
Counts are per test function, not per parameter.
terminalbash
# who owns what, without running anything
tj audit owners -s tests

# in CI: fail when a test has no owner or feature
tj audit owners -s tests --fail-on-gaps owners,features

# which Checkout tests have no tag, by component
tj audit components -s tests --features Checkout --no-tags

↑ Back to the list

13Metadata & Targeting

Execution priority order

Run the smoke checks and the slowest suites first: failures surface sooner, and the long tail doesn't hold up the end of the run.

priority= on suites and tests; lower runs first.
  • Suites with a priority start first, then the rest; non-parallel suites without a priority run last, and the same rules order tests inside a suite.
  • In a parallel run, a prioritized suite keeps its turn until a thread frees up instead of being overtaken.
Limits: no relative ordering ("run after that test") or dependency ordering.
example.pypython
@Suite(priority=1)
class SmokeSuite: ...

@Suite(priority=2)
class CheckoutSuite:

    @test(priority=1)
    def adds_item(self): ...

↑ Back to the list

14Metadata & Targeting

Per-suite test order strategy

Shuffling tests flushes out hidden dependencies between them, but shuffling everything at once makes every run different.

@Suite(order=...) picks ALPHABETICAL, RANDOM, PRIORITY_ASC or PRIORITY_DESC for one suite.
  • --seed N repeats a random order.
  • The header prints the seed whenever a suite is shuffled, and the JSON report records it.
Only tests inside a suite are shuffled, not the order of suites.
example.pypython
from test_junkie.constants import TestOrder

@Suite(order=TestOrder.RANDOM)        # shuffle just this suite
class CartSuite: ...

# tj run --seed 4242                  repeat yesterday's order

↑ Back to the list

15Metadata & Targeting

CLI targeting & run filters

CI runs smoke tests on every push and the full set at night; an owner wants just their tests. Picking tests should take a flag, not a script.

  • --tags-any / --tags-all and their --skip-tags-* pair.
  • --owners, --components and --features, each a flag of its own.
  • --tests (names, Suite.test or patterns) and -x for suites.
The same filters work in Runner.run() (tags through tag_config), and can be saved in the project config.
terminalbash
tj run --tags-any smoke --skip-tags-any slow
tj run --owners payments-team --features Checkout
tj run -t "CheckoutSuite.pays_*"
Limits
Tags combine through the any / all / skip pairs, not a free-form boolean expression.

↑ Back to the list

16Metadata & Targeting

Runtime metadata updates

Some facts are only known while the test runs: which order it created, which known bug it hit.

Meta.update(order_id=order.id) in a running test, any helper it calls, or @beforeTest / @afterTest sets metadata for exactly the parameter, suite parameter and attempt that's running, with no self to pass.
  • Meta.get() reads it back and Meta.bind() carries it into threads the test starts.
  • Meta.link(label, url) and Meta.attach(name, bytes_or_path) add links and files.
  • Listeners get it in properties["test_meta"] (per attempt in properties["meta_attempts"]).
example.pypython
@test(parameters=CARDS)
def charges_card(self, parameter):
    order = shop.checkout(parameter)
    Meta.update(order_id=order.id)
In the reports
The JSON report adds meta per test and meta_set per run; the HTML report shows a metadata panel with clickable links and attachments embedded up to 512 KB; the XML report writes JUnit <properties>. Owner, tags and component can't change at run time.

↑ Back to the list

17Results & Reporting

Clear test outcomes

When login breaks in setup, a report that says "41 failed" sends 41 people looking at the wrong thing. The useful answer is "1 setup error, 40 tests never ran".

Six outcomes, each meaning one thing:
  • fail vs error: an AssertionError is a failure (the product is wrong); any other exception is an error (the test or the environment broke). The same split applies to @beforeClass, @afterClass and group hooks, each with its own listener event.
  • ignored: tests that never got to run because @beforeClass or @beforeGroup failed. They're recorded per parameter, each carrying the error that stopped it, so with two accounts only the account whose login broke is ignored.
  • cancelled: tests a Ctrl+C or Runner.cancel() stopped before they started. Reports are still written.
  • skipped and success.
example.pypython
@Suite(parameters=["admin", "viewer"])
class BillingSuite:

    @beforeClass()
    def login(self, suite_parameter): ...     # raises for "viewer"

    @test(parameters=pages)
    def page_opens(self, parameter, suite_parameter): ...

# admin x pages: run normally
# viewer x pages: ignored, each one carrying the login error
Good to know
Ignored tests count as unsuccessful, so retries and --rerun pick them up.

↑ Back to the list

18Results & Reporting

HTML / XML / JSON reports

People read HTML, CI reads XML, scripts read JSON.

All three ship with Test Junkie: --html-report, --xml-report, --json-report (a folder works too). With -m, CPU and memory are charted in the HTML report and in the console. Screenshots and other files from Meta.attach() are embedded in the HTML report (up to 512 KB each, bigger ones saved next to it). See a sample HTML report.
terminalbash
tj run -m --html-report reports/ --xml-report reports/ --json-report reports/
HTML is a single file to attach to CI or email. It breaks results down by feature, component, owner and tag, and points things out for you: tracebacks shared by several failures, time lost to retries, and on a serial run, how much threading would save, with the slowest tests listed.
JSON records every retry attempt with its status, time and error, plus the parameters and the seed, whether the test was flaky, its metadata and conflict waits. It's what --rerun reads.
XML is JUnit that CI dashboards read: classname and time per test, parameters in the test case name (totals[viewer]), the exception's type, message and <stackTrace>, retried runs recorded the Maven Surefire way (<flakyFailure> / <rerunFailure>), and metadata as <properties>.

↑ Back to the list

19Results & Reporting

Per-suite event listeners

Results belong in your own systems as they happen: a database, Slack, a dashboard.

A Listener class per suite (@Suite(listener=...)), with 22 named events: 9 per test (including on_retry), 9 per suite and 4 for group hooks, including separate events for a @beforeClass failure versus error.
  • Each event gets the live SuiteObject and TestObject and that run's parameters, and the full exception even when reports truncate it.
  • on_failure fires on every attempt and on_complete once with the final result, so one listener can log flakiness and another record the outcome.
example.pypython
class ResultsToDb(Listener):

    def on_complete(self, **kwargs):          # once per test x parameter
        test = kwargs["properties"]["jm"]["jto"]
        account = kwargs["properties"]["suite_meta"]["parameter"]
        db.insert(test=test.get_function_name(), account=account)
Good to know
A listener that raises can't end its suite: the error is collected, and run() raises once reports are written. There's no run-wide listener or run start/end event; share one class across suites.

↑ Back to the list

20Results & Reporting

Programmatic result access

After a run, scripts decide what happens next: open tickets, compare with yesterday, gate a deploy.

Runner.run() returns an aggregator with the run's statistics, already grouped: get_report_by_owner(), get_report_by_features() and get_report_by_tags(), so routing failures to each team is one call. If the run itself errored, it raises after writing reports instead.
  • runner.get_executed_suites() always returns the SuiteObjects and TestObjects, down to each parameter, suite parameter and attempt: status, timing, retries and the policy behind them, metadata per attempt, hook timings, captured output, owners and tags.
  • The exceptions are real exception objects, so isinstance works.
example.pypython
runner = Runner([CheckoutSuite])
runner.run()
for suite in runner.get_executed_suites():
    for test in suite.get_test_objects():
        print(test.get_function_name(), test.get_owner())
Act on the results
A script can build a targeted rerun with Rerun().add(...) and pass it straight back to Runner.run(rerun=...). test.is_flaky() / get_flaky() say which parameters passed only on a retry.

↑ Back to the list