Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
d3e2503
Phase 1 scaffolding: version 1.0.0-DEV, Julia 1.10 floor, design docs
adolgert Sep 27, 2026
f6c8835
Add the 1.0 semantic contract as a developer page
adolgert Sep 27, 2026
62ab8df
Add the independent checker, random-problem gate, and fixture inventory
adolgert Sep 27, 2026
704dab6
Apply Phase 1 review: tighten the gate, fix two contract clauses, fre…
adolgert Sep 27, 2026
eadb756
Add RuleTable, the index-space boundary between the model and feasibi…
adolgert Sep 27, 2026
530c181
Add TestSpace, Invalid and Partition wrappers, constraints, macros, t…
adolgert Sep 27, 2026
67d87b4
Add index-space feasibility: completable, dead, classify, explain_par…
adolgert Sep 27, 2026
3884841
Add isallowed and explain over TestSpace; flip the Phase 2 pending tests
adolgert Sep 27, 2026
98093ec
Apply Phase 2 review: wrapper guard, macro scoping, budget and memo s…
adolgert Sep 27, 2026
2d8df86
Add the internal Request and Design types and move the engine structs
adolgert Sep 27, 2026
234fe9d
IPOG commits values only while the row stays completable; remove disa…
adolgert Sep 27, 2026
b955f50
GND progress guarantee; excursions and full factorial on the request …
adolgert Sep 27, 2026
cbe19c9
Turn on the random gate for both engines; benchmark baseline; close #51
adolgert Sep 27, 2026
6df2fa2
Apply Phase 3 review: single-distance excursions, kept explanations,
adolgert Sep 27, 2026
712e091
Record the CI-runner benchmark baseline and compatibility matrix
adolgert Sep 27, 2026
dff7062
Add TestCases and Exclusion; CI runs Julia 1.10, 1.13, and latest
adolgert Sep 27, 2026
f97954e
Phase 4: the public interface and TestCases
adolgert Sep 27, 2026
2650c6a
Apply Phase 4 review: alias conflicts error, must_include read once,
adolgert Sep 27, 2026
e8b7e1c
Apply Phase 4 review round 2: limit before search, keyword validation,
adolgert Sep 27, 2026
b13bc13
Add coverage and missing_interactions; delete coverage_set.jl
adolgert Sep 27, 2026
858d521
Add report and design_sizes
adolgert Sep 27, 2026
67455c5
Apply Phase 5 review; target version 0.5.0 instead of 1.0.0
adolgert Sep 27, 2026
840ea13
Phase 6: Invalid generation, Partition realization, diagnose, github_…
adolgert Sep 27, 2026
2265038
Apply Phase 6 review: followups search negative rows, negative progress
adolgert Sep 28, 2026
7268faf
Docs: move to Documenter 1 and stop tracking docs/Manifest.toml
adolgert Sep 28, 2026
bbce068
Docs: the Phase 7 page tree (pages to be written)
adolgert Sep 28, 2026
3da13be
Docs: placeholder pages for the Phase 7 tree
adolgert Sep 28, 2026
d6e62c2
Phase 7: documentation and first contact
adolgert Sep 28, 2026
4d9d424
Phase 7 polish: documented Invalid.value, the <: macro message, diagnose
adolgert Sep 28, 2026
7115957
Merge pull request #58 from adolgert/feature/phase7-docs
adolgert Oct 3, 2026
bec4784
Merge pull request #57 from adolgert/feature/phase6-features
adolgert Oct 3, 2026
1858689
Merge pull request #56 from adolgert/feature/phase5-measure
adolgert Oct 3, 2026
0415608
Merge pull request #55 from adolgert/feature/phase4-interface
adolgert Oct 3, 2026
139012d
Merge pull request #54 from adolgert/feature/phase3-engines
adolgert Oct 3, 2026
4df1ad3
Merge pull request #53 from adolgert/feature/phase2-model
adolgert Oct 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 26 additions & 6 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
name: CI
on:
- push
- pull_request
push:
pull_request:
workflow_dispatch:
jobs:
test:
name: Julia ${{ matrix.version }} - ${{ matrix.os }} - ${{ matrix.arch }} - ${{ github.event_name }}
Expand All @@ -10,17 +11,18 @@ jobs:
fail-fast: false
matrix:
version:
- 'lts'
- '1.11'
- '1.10'
- '1.13'
- '1'
os:
- ubuntu-latest
arch:
- x64
include:
- version: "1.11"
- version: "1.13"
os: macOS-latest
arch: arm64
- version: "1.11"
- version: "1.13"
os: windows-latest
arch: x64
steps:
Expand Down Expand Up @@ -61,3 +63,21 @@ jobs:
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # If authenticating with GitHub Actions token
DOCUMENTER_KEY: ${{ secrets.DOCUMENTER_KEY }}
benchmark:
# The Phase 3 performance baseline on a stable runner (design/benchmark_procedure.md).
# GND on the 15x4 strength-4 fixture is left out (--skip-slow); it takes about four minutes per call.
name: Benchmark - ubuntu-latest - Julia 1
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: julia-actions/setup-julia@v2
with:
version: '1'
- uses: julia-actions/cache@v2
- uses: julia-actions/julia-buildpkg@v1
- name: Run the benchmark script
run: julia --color=yes --project benchmark/run.jl --skip-slow --runs 5 --out benchmark_results.md
- uses: actions/upload-artifact@v4
with:
name: benchmark-results-${{ github.sha }}
path: benchmark_results.md
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -2,3 +2,6 @@
*.jl.cov
*.jl.mem
/docs/build/
.DS_Store
Manifest.toml
.vscode/
6 changes: 4 additions & 2 deletions Project.toml
Original file line number Diff line number Diff line change
@@ -1,13 +1,15 @@
name = "UnitTestDesign"
uuid = "239896fa-e45a-40e8-9993-3c434b0bc450"
authors = ["Andrew Dolgert <adolgert@uw.edu>"]
version = "0.4.0"
version = "0.5.0-DEV"

[deps]
Combinatorics = "861a8166-3701-5b0c-9a16-15d98fcdc6aa"
JSON = "682c06a0-de6a-54ab-a142-c8b1cf79cde6"
Random = "9a3f8284-a2c9-5f02-9a11-845980a1fd5c"

[compat]
Combinatorics = "^1"
JSON = "1"
Random = "^1"
julia = "^1.2"
julia = "1.10"
198 changes: 184 additions & 14 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,29 +3,199 @@
[![Stable](https://img.shields.io/badge/docs-stable-blue.svg)](https://adolgert.github.io/UnitTestDesign.jl/stable)
[![Dev](https://img.shields.io/badge/docs-dev-blue.svg)](https://adolgert.github.io/UnitTestDesign.jl/dev)
[![Build Status](https://github.com/adolgert/UnitTestDesign.jl/workflows/CI/badge.svg)](https://github.com/adolgert/UnitTestDesign.jl/actions)
[![Coverage](https://codecov.io/gh/adolgert/UnitTestDesign.jl/branch/master/graph/badge.svg)](https://codecov.io/gh/adolgert/UnitTestDesign.jl)
[![Coverage](https://codecov.io/gh/adolgert/UnitTestDesign.jl/branch/main/graph/badge.svg)](https://codecov.io/gh/adolgert/UnitTestDesign.jl)

Chooses function arguments to make unit testing faster and more effective.
Describe the configurations your code must handle; it tells you which
combinations your tests exercise, and supplies a compact set of additional
cases covering the rest.

* [Documentation](http://computingkitchen.com/UnitTestDesign.jl/stable/)
```
pkg> add UnitTestDesign
```

This package generates parameter values for unit tests, chooses software configurations for integration testing, or generates test datasets. If the system-under-test takes a long time to run or has many possible parameters or many possible values each parameter can take, this library chooses combinations of parameters that are more likely to find faults in the code. It assumes that code will break when there are _interactions_ between different parameter choices, so it generates test data that covers all possible interactions among two parameters, the **all-pairs** algorithm, or three parameters, the **all-triples** algorithm, or higher-order combinatorial interactions.
That installs 0.5 once 0.5 is registered. Until then, install the release
branch with `pkg> add https://github.com/adolgert/UnitTestDesign.jl#release/0.5`.
It needs Julia 1.10 or later.

## When to use it

## Installation
| If you have… | Use… |
|:--|:--|
| Several parameters, a few representative values for each, and bugs that plausibly live in *combinations* of them (an `if` on one option inside a branch on another) | A covering design: `all_pairs`, `all_triples`, or `covering(space; strength)` |
| The same, but each run is cheap and the full product is small | Every combination: `full_factorial`, or `Iterators.product` |
| One known-good configuration, and a question about which single or paired changes break it | `excursions` |
| Hand-written tests already, and a question about which combinations they miss | `coverage`, then `must_include` to add cases for the gaps |
| Classes of input to combine, and values to draw within each class | A covering design over `Partition`s picks each parameter's class; a generator draws the value |
| **Use something else:** values you can generate but not list (strings, trees, arbitrary floats), cheap runs, and a hunt for the one input that breaks a routine | Property-based testing, which explores values and shrinks failures ([Supposition.jl](https://github.com/Seelengrab/Supposition.jl)), or fuzzing |
| **Use something else:** a question of how much each factor affects an outcome | Design of experiments. Orthogonal arrays and fractional factorials are balanced for estimation; covering designs are not. |
| **Use something else:** a sweep over continuous parameters | Space-filling samples, such as Sobol sequences or Latin hypercubes |

```
pkg> add UnitTestDesign
## Example

A solver takes a mode, a factorization and a tolerance. The factorization
applies only in exact mode, and exact mode needs a tight tolerance. Write the
parameters and those two rules as a `TestSpace`, and ask for every pair of
values:

```julia-repl
julia> using UnitTestDesign

julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);

julia> cases = all_pairs(space)
5 cases · strength 2 · IPOG · 3 parameters · 12 combinations
excluded: 3 pairs forbidden, 2 impossible under the constraints; see report(cases)
mode solver tol
1 :exact :qr 1.0e-6
2 :exact :lu 1.0e-6
3 :exact :none 1.0e-6
4 :fast :none 0.001
5 :fast :none 1.0e-6

julia> explain(space, (solver = :lu, tol = 1e-3))
infeasible: no valid case contains (solver = :lu, tol = 0.001); rules 1 and 2 together exclude it (rule 1: @require(mode == :exact || solver == :none); rule 2: exact mode needs a tight tolerance)

julia> handwritten = [(mode = :fast, solver = :none, tol = 1e-3),
(mode = :exact, solver = :lu, tol = 1e-6)];

julia> coverage(handwritten, space)
covers 6 of 11 feasible pairs, 5 missing: (mode = :exact, solver = :none), (mode = :exact, solver = :qr), (mode = :fast, tol = 1.0e-6), (solver = :none, tol = 1.0e-6), (solver = :qr, tol = 1.0e-6)
excluded: 3 pairs forbidden, 2 impossible under the constraints
```

## Example
The five cases hold every pair of values that some valid case can hold. The
excluded line counts the pairs that no valid case can hold, and `explain` names
the rules behind one of them, a pair that neither rule mentions on its own.
`coverage` measures the interaction coverage of tests you already have, and
`all_pairs(space; must_include = handwritten)` keeps those two cases first and
adds cases for the five missing pairs. Each case is a `NamedTuple`, so a test
loops over them:

```julia
test_set = all_pairs(
[1, 2, 3], ["low", "mid" ,"high"], [1.0, 3.7, 4.9], [:greedy, :relax, :optim]
)
for test_case in test_set
test_result = function_under_test(test_case...)
@test test_result == known_result(test_case)
@testset "solve" begin
for (; mode, solver, tol) in cases
@test solve(A, b; mode, solver, tol) ≈ A \ b
end
end
```

The saving grows with the number of parameters: four parameters of three
values each have 81 combinations, and `all_pairs` covers every pair of their
values in 10 cases.

## What it promises

- Every returned case is valid: it satisfies every constraint.
- Every combination the design asks for (every pair of values at strength 2,
every triple at strength 3) that at least one valid case contains appears in
at least one returned case. You never write the rules that other rules
imply.
- Every combination it leaves out is attributed: *forbidden* by rules it
names, or *impossible* because the rules it names combine.
- An unknown is never disguised. If a search reaches its budget, generation
stops with a `ResourceLimitError` that names the limit, and `coverage`
reports the combination as unresolved, with its counts as bounds and no
percentage. Nothing is called covered, excluded or complete that was not
decided.
- The same call, under the same package and Julia versions, gives the same
cases. `IPOG`, the default engine, uses no randomness; `GND` draws from a
fixed default seed. To keep a list of cases across releases and edits,
commit it, or pass it back as `must_include`.

The designs are compact, with no promise of a minimum number of cases;
`design_sizes` shows how many cases each strategy gives before you choose one.
The full statement is the [contract](docs/src/dev/contract.md).

* [Documentation](https://adolgert.github.io/UnitTestDesign.jl/stable), with a
[tutorial](docs/src/man/tutorial.md) that builds up the example above
* [A one-page guide for AI coding agents](docs/src/man/agents.md)
* [JuliaCon 2021 talk](https://www.youtube.com/watch?v=3KIE3yrQ3lw) (YouTube;
it shows 0.4, and the ideas carry over)

## By situation

### A function with many options

Name its parameters, list a few representative values for each, add the rules
for combinations it does not accept, and loop a `@testset` over
`all_pairs(space)`. See [Test a function with many
options](docs/src/howto/many_options.md).

### Generic code across types

Types are values, so a parameter's domain can be `[Int8, UInt64, Float32,
BigInt]`, and a rule can say which element and accumulator types go together.
See [Test generic code across types](docs/src/howto/generic_types.md).

### A CI matrix

`github_matrix(all_pairs(space))` writes the cases as a GitHub Actions
`include:` list, so every pair of operating system, Julia version and option
runs in some job. See [Plan a CI matrix](docs/src/howto/ci_matrix.md).

### A simulation campaign

When each case is a cluster job, write the cases to a table, run them, and
measure what ran with `coverage`. See [Run a simulation
campaign](docs/src/howto/simulation_campaign.md).

### An existing test suite

`coverage(existing, space)` names the combinations your tests miss, and
`all_pairs(space; must_include = existing)` keeps your tests and adds cases for
the gaps. See [Audit and extend an existing
suite](docs/src/howto/audit_existing.md).

### A failure to explain

`diagnose(cases, passed)` ranks the combinations that appear only in failing
cases, and `followups` proposes a case to separate each one. The ranking is a
set of hypotheses, not a proof. See [Diagnose a
failure](docs/src/howto/diagnose.md).

### Invalid inputs

Mark a value `Invalid(x)` and each negative case holds exactly one invalid
value, so one error cannot hide another. See [Test invalid
inputs](docs/src/howto/invalid_inputs.md).

### Property-based testing alongside

A `Partition` names a class of values and draws a concrete value at run time,
so the design chooses the classes and a generator chooses within them. See
[Combine with property-based testing](docs/src/howto/property_based.md).

### A design to commit

Print the cases with `repr(collect(cases))`, paste them into a test file, and
keep two checks: `iscomplete(coverage(CASES, space))` and
`all(case -> isallowed(space, case), CASES)`. The second catches a new rule
that forbids a committed case, which `coverage` lists as rejected rather than
missing. See [Commit a design as data](docs/src/howto/commit_design.md).

## Upgrading from 0.4

0.5 is a breaking release.

- `disallow` is gone. Name the parameters and write the rule as a constraint:
`disallow = (n, level, value, kind) -> level == "high" && kind == :optim`
becomes `all_pairs((n = …, level = …, value = …, kind = …); constraints =
[@forbid(level == "high" && kind == :optim)])`. A rule sees only complete
values, never `nothing` for a parameter not yet chosen.
- Positional calls such as `all_pairs([1, 2], ["a", "b"])` return a
`TestCases{Tuple{...}}`, a read-only vector of tuples, where 0.4 returned a
`Vector{Vector{Any}}`. Loops, destructuring, indexing and `f(case...)` work
as before; code that modifies a row or pushes onto the result does not.
- `n_way`, `seeds`, `wayness`, `all_tuples`, `values_excursion`,
`pairs_excursion`, `triples_excursion` and `GND(M = …)` still work, with a
deprecation warning. Their replacements are `strength`, `must_include`,
`stronger`, `covering`, `excursions(…; distance)` and
`GND(candidates = …)`.
- `generate_tuples`, `Excursion` and the `Counter` keyword are removed.

The [migration table](docs/src/reference/migration.md) lists every change.
5 changes: 5 additions & 0 deletions benchmark/Project.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
[deps]
UnitTestDesign = "239896fa-e45a-40e8-9993-3c434b0bc450"

[sources]
UnitTestDesign = {path = ".."}
12 changes: 12 additions & 0 deletions benchmark/fixtures.jl
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
# The benchmark's fixtures, loaded from the checked-in test definitions so
# that `bench12` has one source of truth (test/fixtures.jl). The fixture
# module needs the checker's input types and the adapter to a TestSpace, as
# the `Checker` test module in test/test_checker.jl does. Nothing here times
# or checks anything.
module BenchFixtures
const TEST = joinpath(dirname(@__DIR__), "test")
include(joinpath(TEST, "checker.jl"))
include(joinpath(TEST, "random_problems.jl"))
include(joinpath(TEST, "fixtures.jl"))
include(joinpath(TEST, "fixture_model.jl"))
end
Loading
Loading