Test Junkie versus pytest, unittest, and Robot Framework — sequential execution, parametrized tests, and parallel execution at 1, 100, and 1,000 tests. Includes package footprint comparison.
Each test body sleeps for exactly 1 ms — enough to make I/O-bound work visible without CPU-saturating the machine.
Wall clock is measured via subprocess.Popen timing in the benchmark harness, making it independent
of each framework's own reporting. Per-phase timing (collection, setup, call, teardown) is collected via each
framework's introspection API where available. Parallel scenarios pass --testlevelsplit
to pabot and -n{N} to pytest-xdist so all frameworks fan out at the individual test level, not suite level.
All results are the median of 100 consecutive runs from a warm filesystem.
Each framework was invoked via subprocess.Popen from the benchmark harness. The commands below are what the harness executed for each scenario. N = test count, W = worker count.
| Framework | Mode | Command |
|---|---|---|
| test_junkie — programmatic API, no CLI; each scenario is a generated Python file run directly | ||
| test_junkie | sequential / parametrized | python bench_N.py — file ends with Runner([BenchSuite]).run() |
| test_junkie | parallel | python bench_N_W.py — file ends with Runner([BenchSuite]).run(test_multithreading_limit=W) |
| test_junkie | collection | python disc_N.py — file instantiates Runner([BenchSuite]) without calling .run() |
| pytest | ||
| pytest | sequential | python -m pytest test_seq_N.py -q --no-header --tb=no |
| pytest | parametrized | python -m pytest test_param_N.py -q --no-header --tb=no |
| pytest | parallel | python -m pytest test_param_N.py -nW -q --no-header --tb=no |
| pytest | collection | python -m pytest test_seq_N.py --collect-only -q --no-header |
| Robot Framework | ||
| Robot Framework | sequential | robot --nostatusrc --quiet --listener=BenchListener.py --outputdir=out suite_seq_N.robot |
| Robot Framework | parametrized | robot --nostatusrc --quiet --listener=BenchListener.py --outputdir=out suite_param_N.robot |
| Robot Framework | parallel | pabot --testlevelsplit --processes=W --nostatusrc --quiet --outputdir=out suite_param_N.robot |
| Robot Framework | collection | robot --dryrun --nostatusrc --quiet --outputdir=out suite_seq_N.robot |
| unittest — programmatic; each scenario is a generated Python file run directly | ||
| unittest | sequential | python test_seq_N.py — file ends with unittest.main(verbosity=0, exit=True) |
| unittest | parametrized | python test_param_N.py — uses @parameterized.expand to generate test methods |
| unittest | parallel | python test_par_N_W.py — file uses ThreadPoolExecutor(max_workers=W) directly; no test runner involved |
| unittest | collection | python disc.py — uses importlib to load the module then TestLoader().loadTestsFromModule(mod) |
Each framework was installed with only what it needs to support the tested feature set: parallel execution and parametrized tests. No extras.
| Framework | Package | Version | Purpose |
|---|---|---|---|
| test_junkie | test-junkie | 0.9a0 | Everything — parallel, parametrization, retries, listeners, reports. No additional packages needed. |
| pytest | |||
| pytest | pytest | 9.1.1 | Core runner, parametrization (@pytest.mark.parametrize) |
| pytest | pytest-xdist | 3.8.0 | Process-based parallel execution (-n N) |
| Robot Framework | |||
| Robot Framework | robotframework | 7.5 | Core runner, parametrization via Test Templates |
| Robot Framework | robotframework-pabot | 5.2.2 | Subprocess-based parallel execution (--testlevelsplit --processes N) |
| unittest | |||
| unittest | unittest | stdlib | Core runner (built in to Python 3.12) |
| unittest | parameterized | 0.9.x | Adds @parameterized.expand — stdlib's subTest() records sub-results but doesn't generate separate test methods, so independent per-variant pass/fail tracking requires this package |
unittest parallel execution is handled via ThreadPoolExecutor with unittest.TestLoader — no package provides a clean test-level parallel runner for unittest, so threading was wired manually in the benchmark harness consistent with how it would be used in practice.
Total installed size including all required plugins to reach feature parity with test_junkie (parallel + parametrization).
| Framework | Core (KB) | Parallel plugin | Param plugin | Total (KB) |
|---|---|---|---|---|
| test_junkie | 299 | built-in | built-in | 299 |
| pytest | 2,970 | +580 (pytest-xdist) | built-in | 3,550 |
| Robot Framework | 5,780 | +510 (robotframework-pabot) | built-in | 6,290 |
| unittest | 250 | built-in (ThreadPoolExecutor) | +118 (parameterized PyPI pkg — stdlib has no @expand-style decorator) |
368 |
Single-threaded execution, one test at a time. Wall clock is the total real-world time from process start to exit, measured externally by the benchmark harness. Overhead = wall clock − (N × 1 ms) — it isolates framework startup, collection, and per-test bookkeeping from the test body itself. Negative overhead means the OS timer resolution made the measured sleep total slightly shorter than N × 1 ms; treat it as zero.
| Framework | N = 1 | overhead | N = 100 | overhead | N = 1,000 | overhead |
|---|---|---|---|---|---|---|
| test_junkie | 145.6 ms | +145 ms | 161.7 ms | +62 ms | 381.6 ms | ~0 ms |
| pytest | 459.0 ms | +458 ms | 679.9 ms | +580 ms | 2,991.6 ms | +1,992 ms |
| Robot Framework | 335.8 ms | +335 ms | 552.3 ms | +452 ms | 2,679.7 ms | +1,680 ms |
| unittest | 165.7 ms | +165 ms | 103.8 ms | +4 ms | 165.4 ms | ~0 ms |
Each framework is invoked once per scenario with all N tests in a single run — the subprocess approach is consistent across all frameworks and matches real-world CLI usage. pytest and Robot Framework are genuinely slower at scale because their per-test internal machinery (hook dispatch, fixture bookkeeping, result recording) carries higher overhead per test than unittest or test_junkie. At N=1,000 that overhead compounds to nearly 2 seconds for pytest. unittest and test_junkie keep overhead near zero regardless of test count.
Single-threaded; N parametrized test variants generated via each framework's native parametrization mechanism.
| Framework | N = 1 | N = 100 | N = 1,000 |
|---|---|---|---|
| test_junkie | 150.2 ms | 149.6 ms | 225.6 ms |
| pytest | 363.2 ms | 692.5 ms | 3,106.9 ms |
| Robot Framework | 332.9 ms | 608.4 ms | 3,260.3 ms |
| unittest | 147.6 ms | 141.7 ms | 204.9 ms |
N parametrized variants distributed across W worker threads/processes. No N=1 row — parallelism only measured at 100 and 1,000 tests.
| Framework | N=100 · W=2 | N=100 · W=5 | N=100 · W=10 | N=1,000 · W=2 | N=1,000 · W=5 | N=1,000 · W=10 |
|---|---|---|---|---|---|---|
| test_junkie | 148.3 ms | 148.8 ms | 227.2 ms | 222.5 ms | 175.4 ms | 171.9 ms |
| pytest | 1,392.4 ms | 2,075.7 ms | 3,308.6 ms | 2,905.9 ms | 3,036.3 ms | 4,275.1 ms |
| Robot Framework | 320.8 ms | 318.0 ms | 316.8 ms | 317.9 ms | 322.0 ms | 318.5 ms |
| unittest | 178.2 ms | 123.8 ms | 108.5 ms | 1,006.5 ms | 464.9 ms | 267.9 ms |
pytest-xdist spawns new interpreter processes per worker — process startup overhead dominates at all test counts, which explains its high parallel wall clock even at N=1,000. Robot Framework's pabot uses subprocesses similarly but has far less per-process overhead, producing consistent ~320 ms regardless of worker count. test_junkie uses threads, so startup cost is near-zero. unittest's ThreadPoolExecutor wins at small N + high W because thread creation is cheap and its per-test overhead is minimal.
Collection = time from session start until tests begin executing. Call = average per-test execution time. Not all phases are measurable for all frameworks via the introspection API used.
| Framework | N | Collection | Avg setup / test | Avg call / test | Avg teardown / test |
|---|---|---|---|---|---|
| test_junkie — collection captured; setup / call / teardown not separately measurable via the Listener API (see note below) | |||||
| test_junkie | 1 | 141.2 ms | — | — | — |
| test_junkie | 100 | 149.9 ms | — | — | — |
| test_junkie | 1,000 | 210.4 ms | — | — | — |
| pytest — full per-phase data via hookwrapper | |||||
| pytest | 1 | 375.4 ms | 0.198 ms | 1.796 ms | 0.175 ms |
| pytest | 100 | 436.8 ms | 0.125 ms | 1.520 ms | 0.134 ms |
| pytest | 1,000 | 554.1 ms | 0.165 ms | 1.754 ms | 0.169 ms |
| Robot Framework — call captured via Listener v3; setup/teardown not separately instrumented | |||||
| Robot Framework | 1 | 327.4 ms | — | 1.877 ms | — |
| Robot Framework | 100 | 387.1 ms | — | 1.764 ms | — |
| Robot Framework | 1,000 | 953.9 ms | — | 1.753 ms | — |
| unittest — collection measured; per-test setup/call hooks limited by API | |||||
| unittest | 1 | 159.0 ms | — | — | <0.01 ms |
| unittest | 100 | 91.6 ms | — | — | <0.01 ms |
| unittest | 1,000 | 96.7 ms | — | — | <0.01 ms |
test_junkie's Listener is an outcome-notification API — it fires events like on_success and on_failure when a test completes, not phase-wrapping hooks that bracket setup, call, and teardown independently. There is no hook that fires exactly at "setup ends / call begins" or "call ends / teardown begins", so the three phases cannot be timed in isolation without modifying the test methods themselves. pytest's hookwrapper mechanism wraps each phase as a separate generator, which is why its numbers are clean. The benchmark attempted to use on_before_test / on_success as approximate call-phase markers, but on_before_test is not a valid Listener event in test_junkie 0.9a0, so the timing was never started and the data came out empty.
Speedup = sequential wall clock ÷ parallel wall clock. Values >1 mean parallel was faster; values <1 mean parallel was slower (startup overhead exceeded gains).
| Framework | N | W=2 | W=5 | W=10 |
|---|---|---|---|---|
| test_junkie | 100 | ×1.01 | ×1.01 | ×0.66 |
| test_junkie | 1,000 | ×1.01 | ×1.29 | ×1.31 |
| pytest | 100 | ×0.50 | ×0.33 | ×0.21 |
| pytest | 1,000 | ×1.07 | ×1.02 | ×0.73 |
| Robot Framework | 100 | ×1.90 | ×1.91 | ×1.92 |
| Robot Framework | 1,000 | ×10.26 | ×10.13 | ×10.24 |
| unittest | 100 | ×0.80 | ×1.14 | ×1.31 |
| unittest | 1,000 | ×0.20 | ×0.44 | ×0.76 |
Robot Framework achieves the highest speedup ratios because its sequential baseline is slow (~3.3 s for 1,000 parametrized tests) — even modest parallel scaling produces impressive-looking multipliers. test_junkie's sequential baseline is already fast enough that parallelism yields near-neutral results for a 1 ms test body; threading overhead roughly cancels the concurrency gain on this hardware. pytest-xdist is net-negative at all N=100 scenarios — process-spawn cost exceeds 3 s per run; at N=1,000 it barely breaks even at W=2 (×1.07) before becoming net-negative again at W=10 (×0.73).
test_junkie & unittest are in the same class
Both complete 1,000 sequential tests in under 400 ms. pytest needs 3 seconds and Robot Framework 2.7 seconds — roughly 8× and 7× slower respectively.
pytest-xdist incurs heavy process startup cost
At 100 tests with W=10 workers, pytest-xdist takes 3.3 s — nearly 5× slower than sequential. test_junkie uses threads, so parallel overhead is near-zero.
test_junkie ships everything; pytest needs 12× more disk
test_junkie is 299 KB, self-contained. Matching its feature set in pytest requires pytest + pytest-xdist = 3,550 KB. Robot Framework + pabot = 6,290 KB.
Collection time dominates at low test counts
For N=1, pytest spends 375 ms collecting and 1.8 ms running the test. test_junkie spends 146 ms total. Startup overhead is the main variable for small suites.