Setting by setting, against pytest, unittest and Robot Framework: what each one does out of the box, what needs a plugin, and how Test Junkie does it, with code. Where another framework does something Test Junkie doesn't, it says so.
Counted from the chips below, one per setting. A chip says whether a capability exists out of the box, not how good it is: read the section.
Execution & Resilience
Suites of UI and API tests spend most of their time waiting on the network. Running them in parallel is the single biggest lever on how long a run takes.
-S and -T): 3 suites can run side by side and share a run-wide pool of 12 test threads. Parallelism is declared where it applies: @Suite(parallelized=False) keeps a whole suite to itself, @test(parallelized=False) runs the test alone: nothing else in the run starts until it and its parameters are done, and parallelized_parameters=True runs a test's parameters side by side (they run one by one by default). Threads share one process, so one logged-in session or one database connection pool can be shared across tests. The trade-off: CPU-heavy tests don't speed up, because threads share one core under the GIL.# tj run -S 3 -T 12 (or Runner.run(suite_multithreading_limit=3, # test_multithreading_limit=12)) @Suite() class CheckoutSuite: @test(parameters=CARDS, parallelized_parameters=True) def pays(self, parameter): ... # cards run side by side # runs alone: nothing else starts until it's done @test(parallelized=False) def resets_inventory(self): ...
Process-based workers (-n auto), with distribution modes such as loadscope, loadfile, loadgroup and worksteal; one worker count, no separate suite and test limits; xdist_group with --dist loadgroup serializes tests by pinning them to one worker. Processes give true multi-core speed for CPU-heavy tests, which threads can't. Each worker process repeats its own session-scoped setup.
Sequential by itself; the third-party unittest-parallel runs TestCase classes or tests in a process pool, with one worker count
Subprocess-based via robotframework-pabot (--processes, default from the CPU count); suite-level by default, test-level with --testlevelsplit; PabotLib for inter-process resource locking
Execution & Resilience
Some tests are fine on their own but break each other: one renames the account another one reads, two suites reset the same feature flag. You want everything in parallel except those pairs.
conflicts_with= on a suite or a test and name what it must never overlap with: suites, tests or a mix, as classes or names ("AccountSettingsSuite.renames_account"), even in other files. That one line is the whole rule: it applies in both directions, so the other side needs no change, and there are no groups, workers or lock names to keep in sync. Everything not named keeps running in parallel, and a suite held back by a conflict doesn't block the queue: the scheduler starts the next suites and comes back to it. Every target is checked before the run starts, and a wrong one stops the run with a link to the docs. tj audit conflicts lists every pair kept apart without running anything, and the console and JSON report (conflict_waits) show how long a test waited and for which test.@Suite(conflicts_with=[AccountSettingsSuite]) # never alongside that suite class LoginSuite: ... @Suite() class ReportsSuite: @test(conflicts_with=[AccountSettingsSuite.renames_account]) def header_shows_account_name(self): ... # never alongside that test
No conflict declaration. @pytest.mark.xdist_group + --dist loadgroup pins a group to one worker, which serializes it. Giving two tests the same group stops them overlapping, but it also serializes everything else in that group: it's placement, not a rule about one pair.
No way to declare a conflict, with or without unittest-parallel
PabotLib Acquire Lock / Release Lock keywords that every test on both sides must call itself; --ordering files with #WAIT barriers, #DEPENDS and { } groups. Nothing checks that a lock call was missed.
Execution & Resilience
UI and integration tests fail for reasons outside your code. Retrying the right things, and only those, keeps a run green without hiding real failures.
@test(retry=N) retries a test; on a parametrized test only the failing combinations run again.@Suite(retry=N) runs the suite again for its unsuccessful tests and repeats @beforeClass / @afterClass only for the suite parameters (e.g. accounts) that still have failures. No other framework re-runs class-level setup selectively like this.tj run --rerun report.json runs only what didn't pass, down to the parameter, and skips setup for accounts with nothing left. It reads an ordinary JSON report, not a local cache: the nightly CI run writes it, and you rerun exactly the failed account × page pairs on your laptop the next morning. From code, Rerun.from_report(path, statuses=[...]) picks which results count.class Flaky(RetryPolicy): when = [...], used as @test(retry=Flaky), @Suite(retry_policy=Flaky) or tj run --retry-policy module:Class: one definition of what to retry, how often and how long to wait, shared by every test that uses it.-p, the console notes the run a test passed on (--flag-flaky labels it FLAKY). Every attempt's traceback is kept, and the HTML report counts retries per test and how much time they cost the run.--retry N sets every test's retry count for one run; --no-retry turns off test and suite retries, e.g. to see the real failure rate. A test that passed only on a retry is flaky: test.is_flaky() and the JSON report's "flaky" say so, tj run --flag-flaky lists them after the summary, and --fail-on-flaky fails the run (exit code 1). The XML report records retried runs the Maven Surefire way (<flakyFailure>, <rerunFailure>), which Jenkins and GitLab read. A retry policy adds delays, backoff and a separate budget per error (see Exception-type retry control).@Suite(parameters=accounts, retry=2) # suite retry: setup only for class NavigationSuite: # accounts that still failed @test(parameters=pages, retry=3) # per account x page pair def page_opens(self, parameter, suite_parameter): ... # tj run --json-report reports/ then tj run --rerun reports/report.json
--reruns N or @pytest.mark.flaky(reruns=N); reruns each failing test item, so only the failing parametrized variants rerun. pytest's built-in --lf / --ff re-run the last failures down to the parametrized item, but from pytest's cache directory, not a shareable report: using them on another machine means copying that directory. --fail-on-flaky fails the run when a test only passed on rerun. Class-scoped setup isn't re-run selectively per parameter. --rerun-show-tracebacks (16.3+) shows failed attempts in the console. JUnit XML records only the final attempt; recording reruns as <flakyFailure> is an open PR (pytest-rerunfailures #380), not released.
Not supported
robotframework-retryfailed adds per-test retries via tags (test:retry(2)) or a global count; --rerunfailed + rebot --merge for a second pass. Built in at keyword level only: Wait Until Keyword Succeeds and WHILE loops. Failed attempts are removed from the log unless the listener is started with RetryFailed:N:True.
Execution & Resilience
A timeout deserves another try. A wrong total on an invoice doesn't. Retrying everything turns real bugs into flaky-looking passes.
retry_on=[...] retries only these exception types, and no_retry_on=[...] never retries these, whatever the retry count. Matching is by class, so subclasses count: retry_on=[requests.exceptions.Timeout] also retries ReadTimeout. When both lists match, no_retry_on wins. The same filters also govern @Suite(retry=) passes, so a suite retry never re-runs a test that failed on a real assertion. For more control, a RetryPolicy also matches the message (When(message="503")). Each condition gets its own budget, delay, backoff and jitter. It can stop retrying after max_time seconds, or for the rest of the run once circuit=N tests failed despite retrying, can run @afterClass + @beforeClass again before a retry (reset="class"), and should_retry() has the last word.from test_junkie.retry import RetryPolicy, When class Flaky(RetryPolicy): when = [When(ConnectionError, attempts=3, delay=1), When(message="503", attempts=2, delay=20, backoff=2)] circuit = 5 # 5 tests still failing: stop @test(retry=Flaky) def refunds_card(self): ...
pytest-rerunfailures: --only-rerun / --rerun-except regexes match "ExceptionName: message" (exception classes too since 16.2); reruns_delay with --reruns-delay-backoff-factor; condition= can inspect the exception (16.7+); --max-suite-reruns caps reruns across the run. One budget and one delay per test, though: no separate attempts or delay per kind of failure. pytest-retry filters by type only with only_on= / exclude= (no message match, no backoff), and flaky takes a rerun_filter function.
No retry in the runner, and no maintained retry plugin. A general library such as tenacity on the test method can filter by type or message, but it retries the method body only: setUp / tearDown don't run again and the runner sees one result
No test-level retry filtered by error, in core or in retryfailed. At keyword level, TRY / EXCEPT with message patterns inside a WHILE loop builds a conditional retry by hand. RF errors are messages, not exception types.
@test(retry=3, retry_on=[ConnectionError, TimeoutError], no_retry_on=[AssertionError]) def charges_card(self): ... # network noise: retry; a wrong answer: fail now
Execution & Resilience
A test that failed and then passed on a retry shows up green in most reports. Whether it's flaky is exactly what you want to know, and CI should be able to act on it.
retry=N, a retry policy or a suite retry pass.test.is_flaky(parameter, suite_parameter) and get_flaky() name each parameter combination that passed only on a retry, with the run it passed on and what failed before. The JSON report adds "flaky" to every test.tj run --flag-flaky lists them after the summary; --fail-on-flaky also fails the run (exit code 1).<flakyFailure> / <flakyError> on a test that passed in the end, <rerunFailure> / <rerunError> on one that never did, each with the exception's type, message and stack trace. Jenkins and GitLab read it.from test_junkie.decorators import Suite, test from test_junkie.retry import RetryPolicy, When class Flaky(RetryPolicy): when = [When(ConnectionError, attempts=3, delay=1), When(message="503", attempts=2, delay=20, backoff=2)] circuit = 5 @Suite(feature="Payments", owner="payments-team") class PaymentsApiSuite: @test(parameters=["visa", "amex"], retry=Flaky) def charges_card(self, parameter): # payments sandbox charge = sandbox.charge(card=parameter, amount=1999) assert charge["status"] == "captured"
pytest-rerunfailures adds --fail-on-flaky, which fails the run when a test passed only on a rerun. JUnit XML records only the final attempt, so CI sees a plain pass; recording reruns as <flakyFailure> is an open PR (pytest-rerunfailures #380), not released.
No retries, so nothing can be flaky in its results
robotframework-retryfailed retries failed tests. By default the failed attempts are dropped from the log, so a flaky pass looks like any other pass; RetryFailed:N:True keeps them.
... Flaky 1 ──────────────────────────────────── PaymentsApiSuite.charges_card [visa] passed on run 2 ConnectionError: Connection reset by peer: sandbox.payments.example FAILED 1 flaky test (--fail-on-flaky) exit code 1
<testcase name="charges_card[visa]" ...> <flakyError type="ConnectionError" message="Connection reset by peer: sandbox.payments.example"> <stackTrace>Traceback (most recent call last): ...</stackTrace> </flakyError> </testcase>
Real output of the suite in this section (the console trimmed to its Flaky block, the XML indented and shortened). Jenkins marks it flaky, not green. A ConnectionError is an error, so it's <flakyError>; a failed assertion would be <flakyFailure>.
Test Design
One test, many inputs: every page, every card brand, every API version, each reported on its own.
parameters=pages() calls the function at import and fixes the list then, like everyone else. parameters=pages passes the function itself, and Test Junkie calls it when the run reaches that test, once per run. A test filtered out by tag or name, or skipped, never builds its matrix, so an API call or database query behind it costs nothing on runs that don't need it. The same goes for suite-level parameter functions: a suite that's filtered out or skipped never calls its function. Any Python object works as a parameter: dicts, dataclasses, page-object classes. Each one is reported by its str(), with no IDs to write, so give your objects a readable __str__, or name them with ids=[...] / ids=lambda p: ... (reports and --rerun use those names). Two parameters that read the same (1 and "1", or objects with the same __str__) are rejected: the test is ignored with an error naming them. Inside a test, skip("reason") skips just the current parameter, and --rerun repeats only the parameters that failed.def pages(): return [Page(DashboardPage), Page(BillingPage, needs="admin")] @test(parameters=pages) # built when the run reaches the test; def page_opens(self, parameter): ... # filtered out or skipped: never built @test(parameters=pages()) # built at import, like everyone else def page_has_title(self, parameter): ...
@pytest.mark.parametrize; mature, accepts any object, with pytest.param(id=, marks=) per value. The parameter list, and so the test count, is fixed at collection (pytest_generate_tests runs then too, before any fixture exists), and every list in the collected files is built before -k or -m deselect anything. Expensive per-value work can be deferred to run time with indirect=True fixtures. Objects get IDs like page0 unless you pass ids=.
self.subTest() records sub-results inside one test; setUp doesn't re-run per variant, there's no decorator, and no retry mechanism at all
Test Templates with inline data rows or FOR loops. robotframework-datadriver generates tests at run time from CSV/Excel or a custom Python reader, with list and dict arguments, at test level only.
Test Design
Many matrices have two layers: who (account, environment, browser) and what (page, endpoint, card). Setup belongs to the first layer, checks to the second.
@Suite(parameters=accounts) × @test(parameters=pages) runs every page for every account. Each pair is tracked, retried and reported on its own. The account reaches @beforeClass(suite_parameter), so you log in once per account rather than once per test. A suite retry repeats setup only for the accounts that still failed, and retry_on / no_retry_on apply per account × page pair. A test that doesn't take suite_parameter runs once, not once per account. Limits: a test's parameter function is built once per run, not per account, and hooks never see the test-level parameter.@Suite(parameters=accounts) class NavigationSuite: @beforeClass() def login(self, suite_parameter): # once per account Browser.login(suite_parameter.email) @test(parameters=pages, parallelized_parameters=True) def page_opens(self, parameter, suite_parameter): parameter.page_object().open() # 2 accounts x 3 pages = 6 runs
@pytest.mark.parametrize on a class applies to every method; stacked with method-level parametrize it produces the full cartesian product. A class-scoped fixture with params=accounts logs in once per account, and pytest reorders tests to keep setups to a minimum.
The third-party parameterized: @parameterized_class makes one class per suite value (so setUpClass runs once per value) and @parameterized.expand multiplies each method; no per-pair retries
No native suite-level × test-level Cartesian product; robotframework-datadriver operates at test level only; all variants must be written or generated explicitly
Test Design
Setup and teardown often depend on the suite's parameter: log in as this account, seed data for this tenant. And some tests in a suite shouldn't run the shared per-test setup at all.
@beforeClass, @afterClass, @beforeTest and @afterTest receive the suite parameter just by declaring suite_parameter; hooks that don't need it leave it out. Within the same suite, any test can opt out of the per-test hooks with skip_before_test / skip_after_test, and out of the Rules hooks with skip_before_test_rule / skip_after_test_rule: one flag on the test, no checks inside the hook. @afterTest still runs when a test is cancelled mid-way. @beforeTest / @afterTest and Rules can also take a test argument: a read-only view of the test they run around (test.get_tags(), test.get_meta(), test.get_owner()), so one hook can branch per test; nothing a hook does to it reaches the run. Hooks receive the suite parameter only, never a test's own parameter, and there's no fixture dependency graph: pytest's fixtures compose and can depend on either parameter level.@Suite(parameters=["admin", "viewer"]) class AccountSuite: @beforeTest() def open_home(self, suite_parameter): # gets the account Browser.open_home(as_user=suite_parameter) @afterTest() def clear_cookies(self): ... # doesn't need it @test(skip_before_test=True) # starts signed out, skips open_home def signed_out_visit_redirects_to_login(self, suite_parameter): ...
The richest model here: fixtures receive parameters through request.param or indirect parametrization, depend on other fixtures, and come in function, class, module, package and session scopes. A single test opts out of an autouse fixture only by overriding the fixture's name or checking a marker inside it.
setUp / setUpClass take no parameters, and subTest variants don't re-run them
Setup keywords take arguments, and [Setup] NONE turns setup off for one test
Metadata & Targeting
When a run of 2,000 tests has 40 failures, the first question is whose they are. Metadata answers it, and drives what runs where.
owner, feature, component, tags and priority are first-class fields, not naming conventions. A wrong type fails when the class loads, naming the argument. They filter runs (--owners, --features, ...), break down the HTML report and tj audit, and reach listeners. One owner= on the suite covers every test, and a test can override it. A team can save its filters in the project config (tj config update), and the run header says when saved config changed a run. meta={...} holds anything else. Not checked: the values themselves, so a misspelled tag goes unnoticed (pytest's --strict-markers catches that).@Suite(feature="Checkout", owner="payments-team") class CheckoutSuite: @test(component="cart", tags=["smoke", "ui"], priority=1, meta={"jira": "PAY-123"}) def adds_item(self): ...
@pytest.mark.X free-form string markers; no first-class owner / component / priority; unregistered markers warn (error under --strict-markers); allure-pytest plugin adds structure
Tests identified by class + method name only; no metadata or tagging system
Built-in [Tags] and suite-wide Test Tags; free name-value Metadata on suites, and on tests since RF 7.5 ([Metadata], with a suite-wide default via Test Metadata). Owner / component / feature / priority are tag strings or metadata by convention (owner:john), with no dedicated filters.
Metadata & Targeting
Tests without an owner rot. An audit shows who owns what and where the gaps are, and a CI gate keeps new tests from arriving without one.
tj audit breaks your tests down by suite, owner, feature, component or tag without running them, with each group's share of the tests audited. --fail-on-gaps fails CI when tests are missing an owner, feature, component or tag (or just the ones you list), and --json feeds dashboards. It combines with the run filters and with gap filters such as --no-owners, --no-tags or --no-test-retries, and the gaps section prints the exact command that lists the offenders. Counts are per test function, not per parameter.# who owns what, without running anything tj audit owners -s tests # in CI: fail when a test has no owner or feature tj audit owners -s tests --fail-on-gaps owners,features # which Checkout tests have no tag, by component tj audit components -s tests --features Checkout --no-tags
--collect-only lists tests; no owner or feature breakdown and no gap gate
Not supported
testdoc (deprecated in RF 7.5, moving to an external tool) writes HTML docs of suites, tests and tags; no grouping by owner and no gap gate
Metadata & Targeting
Run the smoke checks and the slowest suites first: failures surface sooner, and the long tail doesn't hold up the end of the run.
priority= on suites and tests; lower runs first. Suites with a priority start first, then the rest; non-parallel suites without a priority run last, and the same rules order tests inside a suite. In a parallel run, a prioritized suite keeps its turn until a thread frees up instead of being overtaken. No relative ordering ("run after that test") or dependency ordering.@Suite(priority=1) class SmokeSuite: ... @Suite(priority=2) class CheckoutSuite: @test(priority=1) def adds_item(self): ...
pytest-order plugin: @pytest.mark.order(N), relative before= / after= and dependency ordering; no built-in field
Alphabetical order by default; sortTestMethodsUsing for custom sort; no priority field
Tests run in file order and suites alphabetically, with 01__-style name prefixes to force an order; --randomize to shuffle; no priority field
Metadata & Targeting
Shuffling tests flushes out hidden dependencies between them, but shuffling everything at once makes every run different.
@Suite(order=...) picks ALPHABETICAL, RANDOM, PRIORITY_ASC or PRIORITY_DESC for one suite. --seed N repeats a random order; the header prints the seed whenever a suite is shuffled, and the JSON report records it. Only tests inside a suite are shuffled, not the order of suites.from test_junkie.constants import TestOrder @Suite(order=TestOrder.RANDOM) # shuffle just this suite class CartSuite: ... # tj run --seed 4242 repeat yesterday's order
Two separate plugins: pytest-randomly for shuffling modules, classes and tests (repeat with --randomly-seed, switch off with -p no:randomly), pytest-order for explicit ordering; no single per-class strategy setting
One sort function per TestLoader via sortTestMethodsUsing; random or per-class strategies need custom loader code
--randomize all|suites|tests (with an optional seed) applies to the whole run; no per-suite strategy and no priority-based ordering
Metadata & Targeting
CI runs smoke tests on every push and the full set at night; an owner wants just their tests. Picking tests should take a flag, not a script.
--tags-any / --tags-all and their --skip-tags-* pair, --owners, --components, --features, --tests (names, Suite.test or patterns) and -x for suites. The same filters work in Runner.run() (tags through tag_config), and can be saved in the project config. Owner, feature and component are flags of their own; for tags, pytest's -m "smoke and not (ui or slow)" is more expressive than the any / all / skip pairs here.tj run --tags-any smoke --skip-tags-any slow
tj run --owners payments-team --features Checkout
tj run -t "CheckoutSuite.pays_*"
-m and -k take full and / or / not expressions with parentheses; no built-in --owner / --component, so owner-style filtering means encoding it in marker names or writing a plugin
Name selection (python -m unittest Suite.test_name) and -k name-pattern matching (3.7+); no metadata-based filtering
robot --include TAG, --exclude TAG, --test NAME, --suite NAME; AND/OR/NOT tag logic; because owner/component/priority are not structured fields, no dedicated flags exist for them
Metadata & Targeting
Some facts are only known while the test runs: which order it created, which known bug it hit.
Meta.update(order_id=order.id) in a running test, any helper it calls, or @beforeTest / @afterTest sets metadata for exactly the parameter, suite parameter and attempt that's running, with no self to pass. Meta.get() reads it back and Meta.bind() carries it into threads the test starts. Meta.link(label, url) and Meta.attach(name, bytes_or_path) add links and files. Listeners get it in properties["test_meta"] (per attempt in properties["meta_attempts"]); the JSON report adds meta per test and meta_set per run; the HTML report shows a metadata panel with clickable links and attachments embedded up to 512 KB; the XML report writes JUnit <properties>. Owner, tags and component can't change at run time.@test(parameters=CARDS) def charges_card(self, parameter): order = shop.checkout(parameter) Meta.update(order_id=order.id)
Built-in record_property / record_testsuite_property add key-value pairs mid-run, written to JUnit XML and visible to report hooks; markers added at run time don't affect selection (pytest's docs warn record_property breaks validation against the latest JUnit XML schema; record_testsuite_property doesn't)
Not supported
Set Test Metadata (RF 7.5+) and Set Suite Metadata set key-value pairs at run time, shown in the log and report; Set Tags / Remove Tags change tags. Older RF versions have suite-level metadata only.
Results & Reporting
When login breaks in setup, a report that says "41 failed" sends 41 people looking at the wrong thing. The useful answer is "1 setup error, 40 tests never ran".
AssertionError is a failure (the product is wrong); any other exception is an error (the test or the environment broke). The same split applies to @beforeClass, @afterClass and group hooks, each with its own listener event.@beforeClass or @beforeGroup failed. They're recorded per parameter, each carrying the error that stopped it, so with two accounts only the account whose login broke is ignored.Runner.cancel() stopped before they started. Reports are still written.--rerun pick them up.@Suite(parameters=["admin", "viewer"]) class BillingSuite: @beforeClass() def login(self, suite_parameter): ... # raises for "viewer" @test(parameters=pages) def page_opens(self, parameter, suite_parameter): ... # admin x pages: run normally # viewer x pages: ignored, each one carrying the login error
A failing fixture reports each dependent test as ERROR at setup, separate from FAILED, but any exception in the test body is FAILED, assertion or not. No distinct "never ran" or "cancelled" outcome. pytest does have xfail / xpass, which Test Junkie doesn't.
Splits failures (assertions) from errors (other exceptions), and has expected failures. A failing setUpClass becomes one error entry, and the class's tests aren't reported at all.
PASS, FAIL and SKIP only. When a suite setup fails, every test in the suite and its child suites is marked failed.
Results & Reporting
People read HTML, CI reads XML, scripts read JSON.
--html-report, --xml-report, --json-report (a folder works too).--rerun reads.classname and time per test, parameters in the test case name (totals[viewer]), the exception's type, message and <stackTrace>, retried runs recorded the Maven Surefire way (<flakyFailure> / <rerunFailure>), and metadata as <properties>.-m, CPU and memory are charted in the HTML report and in the console. Screenshots and other files from Meta.attach() are embedded in the HTML report (up to 512 KB each, bigger ones saved next to it). There's nothing like Robot's keyword-by-keyword log.html.tj run -m --html-report reports/ --xml-report reports/ --json-report reports/
JUnit XML built in (--junitxml); HTML needs pytest-html; JSON needs pytest-reportlog or pytest-json-report (last release 2022); Allure via plugin; no system resource graphs
TextTestRunner terminal output built in; third-party unittest-xml-reporting (xmlrunner) for XML and html-testRunner for HTML; no JSON
The richest built-in reports here: log.html traces every keyword step, plus report.html and output.xml on every run; XUnit XML via --xunit; JSON output via --output output.json (RF 7.0+); no resource graphs
Results & Reporting
Results belong in your own systems as they happen: a database, Slack, a dashboard.
Listener class per suite (@Suite(listener=...)), with 22 named events: 9 per test (including on_retry), 9 per suite and 4 for group hooks, including separate events for a @beforeClass failure versus error. Each event gets the live SuiteObject and TestObject and that run's parameters, and the full exception even when reports truncate it. on_failure fires on every attempt and on_complete once with the final result, so one listener can log flakiness and another record the outcome. A listener that raises can't end its suite: the error is collected, and run() raises once reports are written. There's no run-wide listener or run start/end event: share one class across suites.class ResultsToDb(Listener): def on_complete(self, **kwargs): # once per test x parameter test = kwargs["properties"]["jm"]["jto"] account = kwargs["properties"]["suite_meta"]["parameter"] db.insert(test=test.get_function_name(), account=account)
Hook functions (pytest_runtest_logreport etc.) in a conftest.py (per directory) or a plugin; not scoped to a test class; you learn the hook protocol
Subclass TestResult (startTest / stopTest, addSuccess / addFailure / addError / addSkip …) and pass it via TextTestRunner(resultclass=...); one result object for the whole run, nothing per class
Listener API in core, the broadest here: suite, test, keyword and control-structure events, with live data and result objects (v3, the default since RF 7.0). Attach run-wide with --listener, or scope it to the suites that import a library via ROBOT_LIBRARY_LISTENER.
Results & Reporting
After a run, scripts decide what happens next: open tickets, compare with yesterday, gate a deploy.
Runner.run() returns an aggregator with the run's statistics, already grouped: get_report_by_owner(), get_report_by_features() and get_report_by_tags(), so routing failures to each team is one call. If the run itself errored, it raises after writing reports instead. runner.get_executed_suites() always returns the SuiteObjects and TestObjects, down to each parameter, suite parameter and attempt: status, timing, retries and the policy behind them, metadata per attempt, hook timings, captured output, owners and tags, and the exceptions are real exception objects, so isinstance works. A script can build a targeted rerun with Rerun().add(...) and pass it straight back to Runner.run(rerun=...). test.is_flaky() / get_flaky() say which parameters passed only on a retry.runner = Runner([CheckoutSuite]) runner.run() for suite in runner.get_executed_suites(): for test in suite.get_test_objects(): print(test.get_function_name(), test.get_owner())
pytest.main() returns an exit code; for per-phase TestReports (setup / call / teardown) you write a small plugin object with hooks and pass it in via pytest.main(plugins=[...]), or read a JUnit/JSON file afterwards. Parameters exist only as part of the test ID (test_x[chrome]), and there's no suite object with totals
TestResult.failures / .errors / .skipped — list of (TestCase, str) tuples; raw traceback strings, no structured exception; per-test timing only via collectedDurations (3.12+)
robot.api.ExecutionResult('output.xml') + ResultVisitor in core; full access to suite/test status, timing and messages, down to each keyword, from saved files too; robot.run() returns a return code; live access during a run via listener v3