diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 2580867..a0d7e33 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -1,7 +1,8 @@ name: CI on: - - push - - pull_request + push: + pull_request: + workflow_dispatch: jobs: test: name: Julia ${{ matrix.version }} - ${{ matrix.os }} - ${{ matrix.arch }} - ${{ github.event_name }} @@ -10,17 +11,18 @@ jobs: fail-fast: false matrix: version: - - 'lts' - - '1.11' + - '1.10' + - '1.13' + - '1' os: - ubuntu-latest arch: - x64 include: - - version: "1.11" + - version: "1.13" os: macOS-latest arch: arm64 - - version: "1.11" + - version: "1.13" os: windows-latest arch: x64 steps: @@ -61,3 +63,21 @@ jobs: env: GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} # If authenticating with GitHub Actions token DOCUMENTER_KEY: ${{ secrets.DOCUMENTER_KEY }} + benchmark: + # The Phase 3 performance baseline on a stable runner (design/benchmark_procedure.md). + # GND on the 15x4 strength-4 fixture is left out (--skip-slow); it takes about four minutes per call. + name: Benchmark - ubuntu-latest - Julia 1 + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + - uses: julia-actions/setup-julia@v2 + with: + version: '1' + - uses: julia-actions/cache@v2 + - uses: julia-actions/julia-buildpkg@v1 + - name: Run the benchmark script + run: julia --color=yes --project benchmark/run.jl --skip-slow --runs 5 --out benchmark_results.md + - uses: actions/upload-artifact@v4 + with: + name: benchmark-results-${{ github.sha }} + path: benchmark_results.md diff --git a/.gitignore b/.gitignore index 3af67b1..a221db5 100644 --- a/.gitignore +++ b/.gitignore @@ -2,3 +2,6 @@ *.jl.cov *.jl.mem /docs/build/ +.DS_Store +Manifest.toml +.vscode/ diff --git a/Project.toml b/Project.toml index 975a843..9dd3606 100644 --- a/Project.toml +++ b/Project.toml @@ -1,13 +1,15 @@ name = "UnitTestDesign" uuid = "239896fa-e45a-40e8-9993-3c434b0bc450" authors = ["Andrew Dolgert "] -version = "0.4.0" +version = "0.5.0-DEV" [deps] Combinatorics = "861a8166-3701-5b0c-9a16-15d98fcdc6aa" +JSON = "682c06a0-de6a-54ab-a142-c8b1cf79cde6" Random = "9a3f8284-a2c9-5f02-9a11-845980a1fd5c" [compat] Combinatorics = "^1" +JSON = "1" Random = "^1" -julia = "^1.2" +julia = "1.10" diff --git a/README.md b/README.md index 9c40d51..d0c585a 100644 --- a/README.md +++ b/README.md @@ -3,29 +3,199 @@ [![Stable](https://img.shields.io/badge/docs-stable-blue.svg)](https://adolgert.github.io/UnitTestDesign.jl/stable) [![Dev](https://img.shields.io/badge/docs-dev-blue.svg)](https://adolgert.github.io/UnitTestDesign.jl/dev) [![Build Status](https://github.com/adolgert/UnitTestDesign.jl/workflows/CI/badge.svg)](https://github.com/adolgert/UnitTestDesign.jl/actions) -[![Coverage](https://codecov.io/gh/adolgert/UnitTestDesign.jl/branch/master/graph/badge.svg)](https://codecov.io/gh/adolgert/UnitTestDesign.jl) +[![Coverage](https://codecov.io/gh/adolgert/UnitTestDesign.jl/branch/main/graph/badge.svg)](https://codecov.io/gh/adolgert/UnitTestDesign.jl) -Chooses function arguments to make unit testing faster and more effective. +Describe the configurations your code must handle; it tells you which +combinations your tests exercise, and supplies a compact set of additional +cases covering the rest. -* [Documentation](http://computingkitchen.com/UnitTestDesign.jl/stable/) +``` +pkg> add UnitTestDesign +``` -This package generates parameter values for unit tests, chooses software configurations for integration testing, or generates test datasets. If the system-under-test takes a long time to run or has many possible parameters or many possible values each parameter can take, this library chooses combinations of parameters that are more likely to find faults in the code. It assumes that code will break when there are _interactions_ between different parameter choices, so it generates test data that covers all possible interactions among two parameters, the **all-pairs** algorithm, or three parameters, the **all-triples** algorithm, or higher-order combinatorial interactions. +That installs 0.5 once 0.5 is registered. Until then, install the release +branch with `pkg> add https://github.com/adolgert/UnitTestDesign.jl#release/0.5`. +It needs Julia 1.10 or later. +## When to use it -## Installation +| If you have… | Use… | +|:--|:--| +| Several parameters, a few representative values for each, and bugs that plausibly live in *combinations* of them (an `if` on one option inside a branch on another) | A covering design: `all_pairs`, `all_triples`, or `covering(space; strength)` | +| The same, but each run is cheap and the full product is small | Every combination: `full_factorial`, or `Iterators.product` | +| One known-good configuration, and a question about which single or paired changes break it | `excursions` | +| Hand-written tests already, and a question about which combinations they miss | `coverage`, then `must_include` to add cases for the gaps | +| Classes of input to combine, and values to draw within each class | A covering design over `Partition`s picks each parameter's class; a generator draws the value | +| **Use something else:** values you can generate but not list (strings, trees, arbitrary floats), cheap runs, and a hunt for the one input that breaks a routine | Property-based testing, which explores values and shrinks failures ([Supposition.jl](https://github.com/Seelengrab/Supposition.jl)), or fuzzing | +| **Use something else:** a question of how much each factor affects an outcome | Design of experiments. Orthogonal arrays and fractional factorials are balanced for estimation; covering designs are not. | +| **Use something else:** a sweep over continuous parameters | Space-filling samples, such as Sobol sequences or Latin hypercubes | -``` -pkg> add UnitTestDesign +## Example + +A solver takes a mode, a factorization and a tolerance. The factorization +applies only in exact mode, and exact mode needs a tight tolerance. Write the +parameters and those two rules as a `TestSpace`, and ask for every pair of +values: + +```julia-repl +julia> using UnitTestDesign + +julia> space = TestSpace( + (mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]); + constraints = [ + @require(mode == :exact || solver == :none), + forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"), + ]); + +julia> cases = all_pairs(space) +5 cases · strength 2 · IPOG · 3 parameters · 12 combinations +excluded: 3 pairs forbidden, 2 impossible under the constraints; see report(cases) + mode solver tol + 1 :exact :qr 1.0e-6 + 2 :exact :lu 1.0e-6 + 3 :exact :none 1.0e-6 + 4 :fast :none 0.001 + 5 :fast :none 1.0e-6 + +julia> explain(space, (solver = :lu, tol = 1e-3)) +infeasible: no valid case contains (solver = :lu, tol = 0.001); rules 1 and 2 together exclude it (rule 1: @require(mode == :exact || solver == :none); rule 2: exact mode needs a tight tolerance) + +julia> handwritten = [(mode = :fast, solver = :none, tol = 1e-3), + (mode = :exact, solver = :lu, tol = 1e-6)]; + +julia> coverage(handwritten, space) +covers 6 of 11 feasible pairs, 5 missing: (mode = :exact, solver = :none), (mode = :exact, solver = :qr), (mode = :fast, tol = 1.0e-6), (solver = :none, tol = 1.0e-6), (solver = :qr, tol = 1.0e-6) +excluded: 3 pairs forbidden, 2 impossible under the constraints ``` -## Example +The five cases hold every pair of values that some valid case can hold. The +excluded line counts the pairs that no valid case can hold, and `explain` names +the rules behind one of them, a pair that neither rule mentions on its own. +`coverage` measures the interaction coverage of tests you already have, and +`all_pairs(space; must_include = handwritten)` keeps those two cases first and +adds cases for the five missing pairs. Each case is a `NamedTuple`, so a test +loops over them: ```julia -test_set = all_pairs( - [1, 2, 3], ["low", "mid" ,"high"], [1.0, 3.7, 4.9], [:greedy, :relax, :optim] - ) -for test_case in test_set - test_result = function_under_test(test_case...) - @test test_result == known_result(test_case) +@testset "solve" begin + for (; mode, solver, tol) in cases + @test solve(A, b; mode, solver, tol) ≈ A \ b + end end ``` + +The saving grows with the number of parameters: four parameters of three +values each have 81 combinations, and `all_pairs` covers every pair of their +values in 10 cases. + +## What it promises + +- Every returned case is valid: it satisfies every constraint. +- Every combination the design asks for (every pair of values at strength 2, + every triple at strength 3) that at least one valid case contains appears in + at least one returned case. You never write the rules that other rules + imply. +- Every combination it leaves out is attributed: *forbidden* by rules it + names, or *impossible* because the rules it names combine. +- An unknown is never disguised. If a search reaches its budget, generation + stops with a `ResourceLimitError` that names the limit, and `coverage` + reports the combination as unresolved, with its counts as bounds and no + percentage. Nothing is called covered, excluded or complete that was not + decided. +- The same call, under the same package and Julia versions, gives the same + cases. `IPOG`, the default engine, uses no randomness; `GND` draws from a + fixed default seed. To keep a list of cases across releases and edits, + commit it, or pass it back as `must_include`. + +The designs are compact, with no promise of a minimum number of cases; +`design_sizes` shows how many cases each strategy gives before you choose one. +The full statement is the [contract](docs/src/dev/contract.md). + +* [Documentation](https://adolgert.github.io/UnitTestDesign.jl/stable), with a + [tutorial](docs/src/man/tutorial.md) that builds up the example above +* [A one-page guide for AI coding agents](docs/src/man/agents.md) +* [JuliaCon 2021 talk](https://www.youtube.com/watch?v=3KIE3yrQ3lw) (YouTube; + it shows 0.4, and the ideas carry over) + +## By situation + +### A function with many options + +Name its parameters, list a few representative values for each, add the rules +for combinations it does not accept, and loop a `@testset` over +`all_pairs(space)`. See [Test a function with many +options](docs/src/howto/many_options.md). + +### Generic code across types + +Types are values, so a parameter's domain can be `[Int8, UInt64, Float32, +BigInt]`, and a rule can say which element and accumulator types go together. +See [Test generic code across types](docs/src/howto/generic_types.md). + +### A CI matrix + +`github_matrix(all_pairs(space))` writes the cases as a GitHub Actions +`include:` list, so every pair of operating system, Julia version and option +runs in some job. See [Plan a CI matrix](docs/src/howto/ci_matrix.md). + +### A simulation campaign + +When each case is a cluster job, write the cases to a table, run them, and +measure what ran with `coverage`. See [Run a simulation +campaign](docs/src/howto/simulation_campaign.md). + +### An existing test suite + +`coverage(existing, space)` names the combinations your tests miss, and +`all_pairs(space; must_include = existing)` keeps your tests and adds cases for +the gaps. See [Audit and extend an existing +suite](docs/src/howto/audit_existing.md). + +### A failure to explain + +`diagnose(cases, passed)` ranks the combinations that appear only in failing +cases, and `followups` proposes a case to separate each one. The ranking is a +set of hypotheses, not a proof. See [Diagnose a +failure](docs/src/howto/diagnose.md). + +### Invalid inputs + +Mark a value `Invalid(x)` and each negative case holds exactly one invalid +value, so one error cannot hide another. See [Test invalid +inputs](docs/src/howto/invalid_inputs.md). + +### Property-based testing alongside + +A `Partition` names a class of values and draws a concrete value at run time, +so the design chooses the classes and a generator chooses within them. See +[Combine with property-based testing](docs/src/howto/property_based.md). + +### A design to commit + +Print the cases with `repr(collect(cases))`, paste them into a test file, and +keep two checks: `iscomplete(coverage(CASES, space))` and +`all(case -> isallowed(space, case), CASES)`. The second catches a new rule +that forbids a committed case, which `coverage` lists as rejected rather than +missing. See [Commit a design as data](docs/src/howto/commit_design.md). + +## Upgrading from 0.4 + +0.5 is a breaking release. + +- `disallow` is gone. Name the parameters and write the rule as a constraint: + `disallow = (n, level, value, kind) -> level == "high" && kind == :optim` + becomes `all_pairs((n = …, level = …, value = …, kind = …); constraints = + [@forbid(level == "high" && kind == :optim)])`. A rule sees only complete + values, never `nothing` for a parameter not yet chosen. +- Positional calls such as `all_pairs([1, 2], ["a", "b"])` return a + `TestCases{Tuple{...}}`, a read-only vector of tuples, where 0.4 returned a + `Vector{Vector{Any}}`. Loops, destructuring, indexing and `f(case...)` work + as before; code that modifies a row or pushes onto the result does not. +- `n_way`, `seeds`, `wayness`, `all_tuples`, `values_excursion`, + `pairs_excursion`, `triples_excursion` and `GND(M = …)` still work, with a + deprecation warning. Their replacements are `strength`, `must_include`, + `stronger`, `covering`, `excursions(…; distance)` and + `GND(candidates = …)`. +- `generate_tuples`, `Excursion` and the `Counter` keyword are removed. + +The [migration table](docs/src/reference/migration.md) lists every change. diff --git a/benchmark/Project.toml b/benchmark/Project.toml new file mode 100644 index 0000000..f5e554f --- /dev/null +++ b/benchmark/Project.toml @@ -0,0 +1,5 @@ +[deps] +UnitTestDesign = "239896fa-e45a-40e8-9993-3c434b0bc450" + +[sources] +UnitTestDesign = {path = ".."} diff --git a/benchmark/fixtures.jl b/benchmark/fixtures.jl new file mode 100644 index 0000000..f02acc8 --- /dev/null +++ b/benchmark/fixtures.jl @@ -0,0 +1,12 @@ +# The benchmark's fixtures, loaded from the checked-in test definitions so +# that `bench12` has one source of truth (test/fixtures.jl). The fixture +# module needs the checker's input types and the adapter to a TestSpace, as +# the `Checker` test module in test/test_checker.jl does. Nothing here times +# or checks anything. +module BenchFixtures + const TEST = joinpath(dirname(@__DIR__), "test") + include(joinpath(TEST, "checker.jl")) + include(joinpath(TEST, "random_problems.jl")) + include(joinpath(TEST, "fixtures.jl")) + include(joinpath(TEST, "fixture_model.jl")) +end diff --git a/benchmark/run.jl b/benchmark/run.jl new file mode 100644 index 0000000..cf89cea --- /dev/null +++ b/benchmark/run.jl @@ -0,0 +1,398 @@ +# The performance baseline of design/benchmark_procedure.md (plan Phase 3 +# step 9). Dependency-free: `@timed` and `Base.summarysize` only. +# +# julia --project=benchmark -e 'using Pkg; Pkg.instantiate()' # once +# julia --project=benchmark benchmark/run.jl [options] +# +# or, with the package's own environment, `julia --project benchmark/run.jl`. +# +# Options: +# --runs N warm runs per measurement (default 5; the median is reported) +# --out PATH also write the Markdown report to PATH +# --tsv PATH also write one tab-separated line per timed measurement +# --skip-slow leave out GND on fixture 1 (about four minutes per call) +# +# Every measurement makes one discarded first call, reported separately as the +# cold start (it includes compilation only for the first measurement that +# reaches a method), then `--runs` warm calls after a `GC.gc()`, and reports +# their median time and median allocation. Case counts are reported so that a +# speedup is never bought with a larger design unnoticed. Tests assert +# completion and coverage, never seconds; this script is for people. +# +# The same script runs on the prior revision (0.4, commit d46122d), which has +# no TestSpace: there it measures fixture 1 only, through the 0.4 API. Point +# its environment at a checkout of that revision, for example +# +# git worktree add ../v04 d46122d +# julia --project=../v04 benchmark/run.jl + +using UnitTestDesign +using InteractiveUtils: versioninfo +using Printf: @sprintf +using Random: Xoshiro + +"True on the 0.4 code, which has no TestSpace and still takes `disallow`." +const LEGACY = !isdefined(UnitTestDesign, :TestSpace) + +LEGACY || include(joinpath(@__DIR__, "fixtures.jl")) + + +## Options + +function options(args) + opts = Dict{String, Any}("runs" => 5, "out" => nothing, "tsv" => nothing, "skip-slow" => false) + i = 1 + while i <= length(args) + a = args[i] + if a == "--skip-slow" + opts["skip-slow"] = true + elseif a in ("--runs", "--out", "--tsv") && i < length(args) + opts[a[3:end]] = a == "--runs" ? parse(Int, args[i + 1]) : args[i + 1] + i += 1 + else + error("unknown option $a; see the header of benchmark/run.jl") + end + i += 1 + end + opts["runs"] >= 1 || error("--runs must be at least 1") + return opts +end + + +## Measuring + +median_of(xs) = (s = sort(collect(xs)); n = length(s); isodd(n) ? s[(n + 1) ÷ 2] : (s[n ÷ 2] + s[n ÷ 2 + 1]) / 2) + +""" +One measurement: `f()` returns `(cases, stats)`, where `cases` is the design's +case count and `stats` the request's `SearchStats` or `nothing`. The first +call is the cold start; then `runs` warm calls. +""" +function measure(f, runs) + GC.gc() + cold = @timed f() + warm = map(1:runs) do _ + GC.gc() + @timed f() + end + cases, stats = last(warm).value + all(w -> first(w.value) == cases, warm) && first(cold.value) == cases || + error("the case count changed between runs of a deterministic call") + return (cases = cases, stats = stats, cold = cold.time, + median = median_of(w.time for w in warm), low = minimum(w.time for w in warm), + high = maximum(w.time for w in warm), bytes = median_of(w.bytes for w in warm), + allocs = median_of(Base.gc_alloc_count(w.gcstats) for w in warm), + gc = median_of(w.gctime for w in warm)) +end + +seconds(t) = t >= 10 ? @sprintf("%.1f", t) : t >= 1 ? @sprintf("%.2f", t) : t >= 0.01 ? @sprintf("%.3f", t) : + @sprintf("%.4f", t) +mib(b) = @sprintf("%.1f", b / 2^20) +count_str(n) = n === nothing ? "–" : replace(string(round(Int, n)), r"(?<=\d)(?=(\d{3})+$)" => ",") + +const TIMING_HEADER = "| Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) |" +const TIMING_ALIGN = "--:|--:|--:|--:|--:|--:|--:|" + +"The header of a table of measurements, with `extra` columns after the timings." +header(first, extra...) = "| $first $TIMING_HEADER" * join(" $x |" for x in extra) * "\n|:--|" * TIMING_ALIGN * + join("--:|" for _ in extra) + +"A table row: the label, the timings, then `extra` cells." +function row(label, m, extra...) + return "| $label | $(m.cases) | $(seconds(m.cold)) | $(seconds(m.median)) | " * + "$(seconds(m.low))–$(seconds(m.high)) | $(mib(m.bytes)) | $(count_str(m.allocs)) | " * + "$(seconds(m.gc)) |" * join(" $x |" for x in extra) +end + +"Queries, nodes and rule checks of a request's feasibility searches." +search_cells(m) = (count_str(m.stats.queries), count_str(m.stats.total_nodes), count_str(m.stats.evaluations)) +const SEARCH_COLUMNS = ("Queries", "Nodes", "Rule checks") + + +## Environment + +function environment(opts) + pkgdir = dirname(dirname(pathof(UnitTestDesign))) + commit = strip(read(Cmd(`git -C $pkgdir rev-parse HEAD`; ignorestatus = true), String)) + dirty = !isempty(strip(read(Cmd(`git -C $pkgdir status --porcelain -- src`; ignorestatus = true), String))) + ci = get(ENV, "CI", "false") == "true" + jl = Base.JLOptions() + lines = [ + "- Date: $(Libc.strftime("%Y-%m-%d %H:%M", time()))", + "- Package: $(pkgdir), commit `$(commit)`$(dirty ? " with uncommitted changes in src/" : "")" * + (LEGACY ? " (0.4 API)" : ""), + "- Julia $(VERSION), $(Sys.MACHINE), $(Sys.KERNEL) $(Sys.ARCH)", + "- CPU: $(Sys.cpu_info()[1].model), $(Sys.CPU_THREADS) logical cores, " * + "$(round(Sys.total_memory() / 2^30, digits = 1)) GiB memory", + "- Julia threads: $(Threads.nthreads()); optimization level -O$(jl.opt_level); bounds checks " * + (jl.check_bounds == 0 ? "default" : jl.check_bounds == 1 ? "on" : "off"), + "- Machine: " * (ci ? "CI runner" : "a workstation or laptop, not a CI runner"), + "- Engine seed: GND `seed = 0` (a fresh `Xoshiro(0)` per call" * (LEGACY ? ", passed as `rng`" : "") * "); IPOG is deterministic", + "- Compilation: the first call of each measurement is reported separately; timings are the median " * + "of $(opts["runs"]) warm calls, each after `GC.gc()`", + ] + return join(lines, "\n") +end + + +## Fixture 1: 15 parameters of 4 values, strength 4, no rules + +fixture1_engine(name) = name == :IPOG ? IPOG() : LEGACY ? GND(rng = Xoshiro(0)) : GND(seed = 0) + +function fixture1(opts, out, tsv) + println(out, "## Fixture 1: 15 parameters × 4 values, strength 4, no rules\n") + println(out, "`all_tuples(fill(1:4, 15)...; n_way = 4, engine)`, the same call on both revisions. " * + "The design fingerprint is `hash` of the returned rows; equal fingerprints mean equal designs.\n") + println(out, header("Engine", "Fingerprint")) + engines = opts["skip-slow"] ? (:IPOG,) : (:IPOG, :GND) + for name in engines + fingerprint = Ref(UInt(0)) + m = measure(opts["runs"]) do + rows = all_tuples(fill(1:4, 15)...; n_way = 4, engine = fixture1_engine(name)) + fingerprint[] = hash(rows) + (length(rows), nothing) + end + println(out, row("$name", m, "`$(string(fingerprint[], base = 16))`")) + println(tsv, join(("fixture1", name, 4, m.cases, m.cold, m.median, m.low, m.high, m.bytes, m.allocs), '\t')) + end + opts["skip-slow"] && println(out, "\nGND left out (`--skip-slow`).") + println(out) +end + + +## Fixture 2: bench12, and fixture 3: the repaired greedy dead ends + +function covering(engine, space, strength) + request = UnitTestDesign.Request(space; strength) + design = UnitTestDesign.generate(engine, request) + return size(design.matrix, 2), request.feasibility.stats +end + +function full_factorial_design(space) + request = UnitTestDesign.Request(space; strength = 1) + design = UnitTestDesign.generate_full_factorial(request) + return size(design.matrix, 2), request.feasibility.stats +end + +function fixture2(opts, out, tsv) + space = BenchFixtures.test_space(BenchFixtures.bench12) + println(out, "## Fixture 2: `bench12` (test/fixtures.jl)\n") + println(out, "12 parameters, 331776 rows, 207360 valid, four rules. `generate(engine, Request(space; strength))` " * + "through the internal request; the public `covering` arrives in Phase 4. Queries, nodes and rule " * + "checks are the request's `feasibility.stats` for one call. Full factorial is " * + "`generate_full_factorial(Request(space; strength = 1))`.\n") + println(out, header("Strategy", SEARCH_COLUMNS...)) + for strength in (2, 3), (name, engine) in ((:IPOG, IPOG()), (:GND, GND(seed = 0))) + m = measure(() -> covering(engine, space, strength), opts["runs"]) + println(out, row("$name, strength $strength", m, search_cells(m)...)) + println(tsv, join(("bench12", name, strength, m.cases, m.cold, m.median, m.low, m.high, m.bytes, m.allocs), '\t')) + end + m = measure(() -> full_factorial_design(space), opts["runs"]) + println(out, row("full factorial", m, search_cells(m)...)) + println(tsv, join(("bench12", :FullFactorial, 0, m.cases, m.cold, m.median, m.low, m.high, m.bytes, m.allocs), '\t')) + println(out) +end + +function fixture3(opts, out, tsv) + println(out, "## Fixture 3: the repaired greedy dead ends\n") + println(out, "The four dead ends frozen in test/fixtures.jl, on which the 0.4 IPOG threw a `BoundsError`. " * + "No \"before\" number exists; these record the \"after\".\n") + println(out, header("Fixture", SEARCH_COLUMNS...)) + for f in (BenchFixtures.dead_end_pairwise_1, BenchFixtures.dead_end_pairwise_2, + BenchFixtures.dead_end_threeway_1, BenchFixtures.dead_end_threeway_2) + space = BenchFixtures.test_space(f) + strength = f.request.strength + for (name, engine) in ((:IPOG, IPOG()), (:GND, GND(seed = 0))) + m = measure(() -> covering(engine, space, strength), opts["runs"]) + println(out, row("`$(f.name)`, $name, strength $strength", m, search_cells(m)...)) + println(tsv, join((f.name, name, strength, m.cases, m.cold, m.median, m.low, m.high, m.bytes, m.allocs), '\t')) + end + end + println(out) +end + + +## The memo of a lazy rule (contract §3.5, §12.19) + +"A copy of `space` with one more rule, a whole-case rule that forbids nothing, so it is lazy." +function with_whole_case_rule(space) + return TestSpace((n => v for (n, v) in zip(space.names, space.values))...; + constraints = [space.constraints; forbid(case -> false)]) +end + +function memo_steps(out, label, space, steps) + size0 = Base.summarysize(space) + println(out, "| $label | before generation | – | – | $(count_str(size0)) | – | – | – |") + for (name, engine, strength) in steps + request = UnitTestDesign.Request(space; strength) + t = @timed UnitTestDesign.generate(engine, request) + println(out, "| $label | after $name, strength $strength | $(size(t.value.matrix, 2)) | " * + "$(seconds(t.time)) | $(count_str(Base.summarysize(space))) | " * + "$(count_str(UnitTestDesign.memo_size(request))) | " * + "$(count_str(Base.summarysize(request.feasibility))) | " * + "$(count_str(request.feasibility.stats.total_nodes)) |") + end +end + +function memo(opts, out) + println(out, "## Memory of a lazy rule's memo\n") + println(out, "Each space gets one added whole-case rule, `forbid(case -> false)`, which is always lazy " * + "(contract §12.19), so every complete row the searches reach is memoized. The memo belongs " * + "to the request (§3.5): `memo_size(request)` counts its entries and " * + "`Base.summarysize(request.feasibility)` is the request's search state, memo included, while " * + "`Base.summarysize(space)` should not change. One space per table; the calls run in order on " * + "the same space, each with a fresh request. Times are single first calls, not medians.\n") + println(out, "| Space | When | Cases | Time (s) | `Base.summarysize(space)` (bytes) | " * + "`memo_size(request)` (entries) | `Base.summarysize(request.feasibility)` (bytes) | Nodes |") + println(out, "|:--|:--|--:|--:|--:|--:|--:|--:|") + bench = with_whole_case_rule(BenchFixtures.test_space(BenchFixtures.bench12)) + memo_steps(out, "bench12 + whole-case rule", bench, + [(:IPOG, IPOG(), 2), (:IPOG, IPOG(), 3), (:GND, GND(seed = 0), 2), (:GND, GND(seed = 0), 3)]) + wide = with_whole_case_rule(TestSpace((Symbol(:p, i) => 1:4 for i in 1:15)...)) + memo_steps(out, "fixture 1 + whole-case rule", wide, [(:IPOG, IPOG(), 2), (:IPOG, IPOG(), 3), (:IPOG, IPOG(), 4)]) + println(out, "\nGND is not run on fixture 1 with the whole-case rule: without rules it already takes about " * + "four minutes per call at strength 4, and the rule adds a feasibility search to every value it scores.\n") +end + + +## Where the time goes (Phase 4/5 candidates; measured, not optimized) + +"Median seconds and bytes of `f()` over `n` calls, after one discarded call." +function quick(f, n = 5) + f() + ts = [(GC.gc(); @timed f()) for _ in 1:n] + return median_of(t.time for t in ts), median_of(t.bytes for t in ts) +end + +function breakdown(opts, out) + println(out, "## Where the time goes\n") + println(out, "Parts of the calls above, timed alone (median of 5 after one discarded call). " * + "Recorded as Phase 4/5 candidates. Phase 3 review round 1 moved the final validation to " * + "index space; Phase 5 made an unconstrained request's target list lazy (`TargetList`) and " * + "its recount streamed, one support at a time.\n") + println(out, "| Call | Part | Time (s) | Allocated (MiB) | Note |") + println(out, "|:--|:--|--:|--:|:--|") + + # Full factorial on bench12: enumeration (each candidate row through + # `violates`) versus the final validation (each accepted row through + # `violates` in index space, since Phase 3 review round 1). + space = BenchFixtures.test_space(BenchFixtures.bench12) + total, total_b = quick(() -> full_factorial_design(space)) + request = UnitTestDesign.Request(space; strength = 1) + design = UnitTestDesign.generate_full_factorial(request) + enumerate_t, enumerate_b = quick() do + r = UnitTestDesign.Request(space; strength = 1) + UnitTestDesign._accept_rows!(Int[], r, Set{Vector{Int}}()) + end + validate_t, validate_b = quick(() -> UnitTestDesign.validate_design(request, design.matrix, Vector{Int}[]; + strategy = :full_factorial)) + # One `forbids` call on a tabulated two-parameter rule, the kind every check here makes. + table = space.tables[1] + key = ones(Int, 12) + calls = 10^6 + forbids_t, forbids_b = quick() do + hits = 0 + for _ in 1:calls + hits += UnitTestDesign.forbids(table, key) + end + hits + end + per_call, per_call_b = forbids_t / calls, forbids_b / calls + r = UnitTestDesign.Request(space; strength = 1) + UnitTestDesign._accept_rows!(Int[], r, Set{Vector{Int}}()) + checks = r.feasibility.stats.evaluations + rows = size(design.matrix, 2) + validate_checks = rows * length(space.tables) + share(t) = @sprintf("%.0f%%", 100 * t / total) + println(out, "| bench12 full factorial | whole call | $(seconds(total)) | $(mib(total_b)) | $(rows) rows of 331776 |") + println(out, "| bench12 full factorial | enumeration, `violates` per candidate | $(seconds(enumerate_t)) | " * + "$(mib(enumerate_b)) | $(share(enumerate_t)) of the call; $(count_str(checks)) rule checks |") + println(out, "| bench12 full factorial | `validate_design`, `violates` per row | $(seconds(validate_t)) | " * + "$(mib(validate_b)) | $(share(validate_t)) of the call; $(count_str(validate_checks)) rule checks |") + println(out, "| `forbids` on a tabulated rule | one call | $(@sprintf("%.1f", 1e9 * per_call)) ns | " * + "$(@sprintf("%.0f", per_call_b)) bytes | the runtime-length `ntuple` key allocates on every check |") + println(out, "| bench12 full factorial | `forbids` in enumeration, estimated | " * + "$(seconds(checks * per_call)) | $(mib(checks * per_call_b)) | $(share(checks * per_call)) of the call |") + println(out, "| bench12 full factorial | `forbids` in validation, estimated | " * + "$(seconds(validate_checks * per_call)) | $(mib(validate_checks * per_call_b)) | " * + "$(share(validate_checks * per_call)) of the call |") + # The round trip through values that validation made before Phase 3 + # review round 1, for comparison: each row becomes a case (`from_indices`, + # which `to_cases` still does), and `isallowed` reads it back + # (`case_indices`, a lookup of each value by identity) before its rule + # checks. Neither is part of the call any more. + as_case(j) = UnitTestDesign.from_indices(space, UnitTestDesign._space_indices(request, design.matrix[:, j])) + cases_t, cases_b = quick(() -> [as_case(j) for j in axes(design.matrix, 2)]) + cases = [as_case(j) for j in axes(design.matrix, 2)] + allowed_t, allowed_b = quick(() -> count(c -> isallowed(space, c), cases)) + lookup_t, lookup_b = quick(() -> sum(c -> sum(UnitTestDesign.case_indices(space, c)), cases)) + println(out, "| bench12 full factorial | not in the call: each row to a case, `from_indices` (`to_cases`) | " * + "$(seconds(cases_t)) | $(mib(cases_b)) | $(share(cases_t)) of the call |") + println(out, "| bench12 full factorial | not in the call: `isallowed` on each case (the old validation) | " * + "$(seconds(allowed_t)) | $(mib(allowed_b)) | $(share(allowed_t)) of the call |") + println(out, "| bench12 full factorial | ... of which `case_indices`, the value lookup | " * + "$(seconds(lookup_t)) | $(mib(lookup_b)) | $(share(lookup_t)) of the call |") + + # Fixture 1 through IPOG: the classic unconstrained `ipog` is the 0.4 + # core; the request adds the lazy target list, which an unconstrained + # request never materializes, and the final validation, which recounts + # it one support at a time. + wide = TestSpace((Symbol(:p, i) => 1:4 for i in 1:15)...) + whole, whole_b = quick(() -> covering(IPOG(), wide, 4)) + targets_t, targets_b = quick(() -> UnitTestDesign.classify_targets(UnitTestDesign.Request(wide; strength = 4))) + core_t, core_b = quick(() -> UnitTestDesign.ipog(fill(4, 15), 4)) + wide_request = UnitTestDesign.Request(wide; strength = 4) + required, _ = UnitTestDesign.classify_targets(wide_request) + matrix = UnitTestDesign.ipog(fill(4, 15), 4) + check_t, check_b = quick(() -> UnitTestDesign.validate_design(wide_request, matrix, required)) + wshare(t) = @sprintf("%.0f%%", 100 * t / whole) + println(out, "| fixture 1, IPOG | whole `generate` call | $(seconds(whole)) | $(mib(whole_b)) | |") + println(out, "| fixture 1, IPOG | classic `ipog`, the 0.4 core | $(seconds(core_t)) | $(mib(core_b)) | $(wshare(core_t)) |") + println(out, "| fixture 1, IPOG | `classify_targets`: the lazy `TargetList` of $(count_str(length(required))) " * + "targets, not materialized | $(seconds(targets_t)) | $(mib(targets_b)) | $(wshare(targets_t)); " * + "unconstrained, so nothing is excluded |") + println(out, "| fixture 1, IPOG | `validate_design`, the recount streamed per support | $(seconds(check_t)) | " * + "$(mib(check_b)) | $(wshare(check_t)) |") + println(out) +end + + +## Main + +function main(args) + opts = options(args) + tsv = opts["tsv"] === nothing ? devnull : open(opts["tsv"], "w") + report = IOBuffer() + # Each section goes to stdout as soon as it is done, and to the report. + function emit(section) + io = IOBuffer() + section(io) + text = String(take!(io)) + print(stdout, text) + flush(stdout) + print(report, text) + end + emit() do io + println(io, "# UnitTestDesign benchmark\n") + println(io, environment(opts), "\n") + println(io, "
versioninfo()\n\n```\n", sprint(versioninfo), "```\n
\n") + end + emit(io -> fixture1(opts, io, tsv)) + if LEGACY + emit() do io + println(io, "Fixtures 2 and 3 are not run on the 0.4 code: its IPOG throws a `BoundsError` on them " * + "and its GND does not return on `bench12` (design/benchmark_procedure.md), so they " * + "have no baseline.\n") + end + else + emit(io -> fixture2(opts, io, tsv)) + emit(io -> fixture3(opts, io, tsv)) + emit(io -> memo(opts, io)) + emit(io -> breakdown(opts, io)) + end + opts["out"] === nothing || write(opts["out"], String(take!(report))) + tsv === devnull || close(tsv) + return nothing +end + +main(ARGS) diff --git a/design/20260926_answers.md b/design/20260926_answers.md new file mode 100644 index 0000000..c94c7b4 --- /dev/null +++ b/design/20260926_answers.md @@ -0,0 +1,7 @@ +The disallow function can be deleted now. No need to support it. +I like the idea of having a function that helps you diagnose outcomes. That could be handy. +I don't know that we need to find evidence for combinatorial testing. That seems like it's a paper to write. Let's focus on that after we get an update done. +We should do Option 3. From Option 4, cherry pick that we should do diagnose, Invalid, Partition, and github_matrix. Those are my favorites. +Otherwise, I like your guesses on how to answer the main questions. Go with it. +Let's be real. You talk about how many months this would take, but we can execute this change quite quickly using AI agents. I'd like you to plan this as a single updated release. Make a plan that has phases, where each phase has steps within it. I'll review at the end of each phase. There should be no more than 8 phases. + diff --git a/design/20260926_implementation_plan.md b/design/20260926_implementation_plan.md new file mode 100644 index 0000000..2a8330c --- /dev/null +++ b/design/20260926_implementation_plan.md @@ -0,0 +1,712 @@ +# UnitTestDesign.jl 1.0: Implementation Plan + +2026-09-27: the release is 0.5.0, not 1.0.0; the maintainer wants to use a +0.x version before calling it 1.0. Read "1.0" below as "0.5". The release +branch is `release/0.5`. + +Date: 2026-09-26. Prepared from `interface_synthesis.md` and the decisions +in `20260926_answers.md`. One release, eight phases, a review at the end of +each phase. Every phase ends with the test suite green and the package +usable, so a review can run the code, not just read it. + +Revised after the plan review: correctness and wrapper semantics are fixed +in Phase 1; implementation remains in eight phases and one release. + +## Decisions carried into this plan + +From your answers: + +- **Option 3 (the friendly product), plus four picks from Option 4:** + `diagnose`, `Invalid`, `Partition`, `github_matrix`. Nothing else from + Option 4: no runner, no outcome files, no `as_code`, no TOML, no + rolling coverage, no `max_cases`. +- **`disallow` is deleted**, not deprecated. `nothing` never denotes an + unassigned parameter in user cases; it remains a legitimate domain value. +- **No evidence track.** The case study and mutation analysis wait until + after this release. +- **The synthesis's leans decide every open question** (its Section + "Questions the best version must answer"). The ones that shape code: + - Constraints live in the `TestSpace`; generation calls accept a + `constraints =` convenience that builds a space. + - All rule forms compile to one internal `Constraint` (scope, predicate, + polarity, label), tabulated over the scope. Whole-case predicates are + allowed as an escape hatch with a documented cost. + - Exclusions are attributed honestly: a named rule for direct + exclusions, a deletion search for implied ones, `unknown` when a search + limit is hit. Never an invented cause. + - `TestCases{T} <: AbstractVector{T}` preserves domain values and their + concrete types without numeric promotion. Homogeneous domains have + concrete field types; heterogeneous domains may use union or abstract + field types. Since this is 1.0, the positional path changes its return type too + (`Vector{Vector{Any}}` becomes `TestCases{Tuple{...}}`). + - `show` computes only bookkeeping; verification and curves live in + `report` and `coverage`. + - Deterministic across runs (GND gets a fixed default seed); not + promised across versions; stable across edits only through + `must_include`. + - Outcomes enter only through the pure function `diagnose(cases, passed)`. + - Julia target is the latest LTS, 1.10, and nothing earlier: `julia = "1.10"` + in `Project.toml`, CI runs on 1.10 (LTS), 1.13, and the latest release, and + the code uses 1.10 features freely (package extensions, `Returns`, + `@NamedTuple`) with no compatibility shims for older versions. + - Vocabulary: `TestSpace`, `constraints`, `must_include` (alias `seeds`), + `strength` (alias `n_way`), `stronger`, `TestCases`, `covering` as the + general entry point (alias `all_tuples`), `coverage`, `explain`, + `report`, `design_sizes`. + +Small choices I made where the synthesis left two names or was silent. +Say so at the Phase 1 review if you want a different one: + +| Choice | Pick | Why | +|:--|:--|:--| +| Planning function | `design_sizes` | Fable-led option; the name says what the table contains | +| Realizing a `Partition` at run time | `realize(case; rng)` | Draws are explicit and seeded by the caller, never hidden in iteration | +| Constraints in positional calls | Not supported; use a space | Rules need names; positional users get every other feature | +| `Invalid` values and rules | In a row whose invalid parameter is `p`, rules mentioning `p` are not applied; all other rules still apply | Contract and coverage meaning reviewed in Phase 1; implemented in Phase 6 | +| The 0.4 `seeds`/`n_way`/`wayness`/`all_tuples`/`*_excursion`/`GND(M=)` spellings | Kept one version as deprecated aliases that warn | Cheap, and it keeps registered users compiling | +| `generate_tuples`, `Excursion`, `Counter` | Removed from the public surface | No user-facing purpose once the request is internal | + +Branching: merge `fix/gnd-match-condition` to `main` first (its commit +message already says only what it fixes). Then one branch, `release/0.5`, +with one PR per phase into it, and one PR from it to `main` in Phase 8. + +## The target, in one screen + +```julia +using UnitTestDesign + +space = TestSpace( + (mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]); + constraints = [ + @require(mode == :exact || solver == :none), + forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"), + ]) + +cases = all_pairs(space) # TestCases{NamedTuple}, 5 cases +explain(space, (solver = :lu, tol = 1e-3)) # infeasible: the two rules together +coverage(handwritten, space) # what an existing suite misses +all_pairs(space; must_include = handwritten) # keep them, add a compact set covering the gaps +report(cases) # excluded list, bonus coverage, prefix curve +design_sizes(space) # cases per strength, before committing +diagnose(cases, passed) # ranked suspects after a run (experimental) +github_matrix(cases) # JSON for a workflow's include: list +all_pairs([1, 2, 3], ["a", "b"], [1.0, 2.0]) # positional still works; returns TestCases{Tuple} +``` + +Final layout of `src/` (new files marked): + +``` +UnitTestDesign.jl exports, includes +space.jl * TestSpace, value identity, Partition and Invalid wrappers +constraints.jl * Constraint, forbid/require, @forbid/@require, tabulation +feasibility.jl * violates, completable (with witness and limit), classify targets +request.jl * internal request: arity, strength, groups, dead(), must_include +combinations.jl (kept) +coverage_matrix.jl (kept; dead-partial predicate replaces disallow) +parameter_order.jl IPOG (kept; three sites use the predicate) +greedy_tuples.jl GND (kept; progress guarantee, seed, candidates) +excursions.jl explicit base, distance +full_factorial.jl incremental enumeration, materialized result, size guard +testcases.jl * TestCases{T}, show, Tables-compatible iteration +interface.jl * covering, all_pairs..., excursions, full_factorial, deprecations +measure.jl * coverage, missing_interactions, report, design_sizes +invalid.jl * one-invalid-per-case generation +partition.jl * realize +diagnose.jl * diagnose, followups +export.jl * github_matrix +``` + +`coverage_set.jl` goes away: the independent checker in `test/` takes over +its role, and `measure.jl` works in value space. + +--- + +## Phase 1: Specification, checker, and scaffolding + +Goal: write down the contract everything else is judged against, and put +the tests in place that will fail until the engines keep it. + +Steps: + +1. **Merge the GND branch to `main`; cut `release/0.5`.** Open one issue, + "Implicit constraints crash IPOG and hang GND," with Opus's + os/gpu/driver example, so Phase 3 has something to close. +2. **Write `docs/src/dev/contract.md`.** The six-point semantic contract + from the synthesis (valid case; feasible combination; every returned + case valid and every required combination covered; exclusions reported + and attributed; user code never sees a partial case; `unknown` under + limits). Add value identity (concrete type and `isequal`, duplicates + rejected, singleton domains allowed, `nothing` is a value), the + determinism boundary, must-include ordering, `stronger` validation + rules, and the "not now" list (runner, outcome files, model inference, + CLI, serialization framework, solver, shrinking, fixture catalog, + `as_code`, TOML, `max_cases`). This page is the spec; later phases cite it. + Resolve these details here, before engine work: + + - **Resource limits.** Feasibility has three states. Generation must + resolve every target and placement decision or stop with a clear + resource-limit error; it never returns an uncertified design. + Measurement and explanation may return `unknown`, and must not claim + complete coverage or an exact percentage with unresolved targets. + Document the `feasibility_limit` keyword, its default node budget, + accounting across component searches, and how to retry with a larger + budget. Explanation limits are separate: proven infeasibility remains + proven even if finding a smaller explanation exhausts its budget. + - **Identity.** Preserve each supplied value and its concrete type in + storage, rules, results, and coverage keys. For example, `Any[1, 1.0]` + contains two choices. No promotion may merge them. Identity of + `Invalid(x)` includes the marker and the wrapped value's type and + `isequal`; `Invalid(x)` and `x` are distinct choices. + - **Partitions.** Within a parameter, partition names are unique; + reject a raw Symbol equal to a partition name to avoid ambiguous rule + patterns. Rules and coverage see the name; returned cases retain the + wrapper until explicit `realize`. Nested wrappers are unsupported. + Realization preserves the tuple or named-tuple shape and makes no + additional coverage claim about the sampled concrete values. + - **Negative cases.** Each parameter must have at least one ordinary + value (a `Partition` counts); reject domains containing only `Invalid` + values. Ordinary rows satisfy every rule. A negative row contains + exactly one `Invalid` value at parameter `p` and satisfies all rules + whose scopes omit `p`. Whole-case rules therefore do not apply to + negative rows. Only ordinary rows contribute to ordinary coverage; + report negative coverage separately. Rows violating either applicable + rule set contribute no coverage. + - **Negative targets.** For each invalid value at `p`, each requested + group containing `p` with strength `s` requires every feasible + `(s-1)`-way combination on the other parameters in that group. + Feasibility uses the negative-row rules above and ordinary values in + every other parameter. Union these targets across the base and + stronger groups. At strength 1, the empty combination requires one + valid completion per invalid value, if one exists. Report impossible + negative targets and stop generation on unresolved ones. Groups + omitting `p` add no negative targets; their ordinary coverage is + already required of the ordinary design. + - **Strategy boundaries.** Full factorial enumerates ordinary and + single-invalid rows satisfying their applicable rules; excursions + use the same row-validity policy within their distance bound. Neither + includes multiple-invalid rows. Must-include rows use the same policy; + partial rows must admit a completion, and keep assigned values intact. + A partial row without an `Invalid` marker is completed as an ordinary + row; negative completion must be requested explicitly with the marker. + - **Honest size claims.** Engines produce compact designs, with no + minimum-case-count guarantee. Deletion search produces a sufficient + explanation, labeled inclusion-minimal only when verified; it does + not promise a minimum-size rule set. +3. **Write the independent checker, `test/checker.jl`.** Value-space, no + shared code with the engines. `check_design(cases, space; strength, + stronger)` enumerates the full product for small spaces, computes the + valid set and the feasible combinations by brute force, and returns + separate ordinary and negative results with valid rows, missing targets, + and rejected rows. Implement the Phase 1 policies independently of + production feasibility and tabulation. Verify the oracle itself against + hand-enumerated fixtures using plain test data in Phase 1; add an adapter + for the production space type in Phase 2. It is the oracle for Phases 3–6. +4. **Write the random problem generator, `test/random_problems.jl`.** + Opus's shape: 3–8 parameters, 2–4 values, 1–4 scoped rules over 2–3 + parameters, half of them written with `!=` or `<`. A `@testitem` runs + 500 pairwise and 500 three-way problems through IPOG and GND, scaled by + `test_run_multiplier()`, with a fixed seed and `seed_mod()`. Record + a narrowly scoped expected failure for the known legacy engine defect; + keep future-API tests explicitly pending until their phase. Do not mask + arbitrary exceptions as expected failures. The random gate is mandatory + in Phase 3. Add a deterministic fixture inventory for disconnected + unsatisfiable components, exhausted limits, heterogeneous values, + partial seeds, overlapping stronger groups, and wrapper interactions. +5. **Bump `Project.toml`** to `1.0.0-DEV`, `julia = "1.10"` (the latest LTS; + nothing earlier is supported). Set the CI matrix to `'1.10'`, `'1.13'`, + and `'1'`, replacing the `lts` alias so the floor is explicit. Add `.DS_Store`, `Manifest.toml`, `.vscode/` + to `.gitignore`. Move `interface_*.{md,pdf,tex}`, `interface_synthesis.*`, + `z3_example.jl`, and the two dated files into `design/` so the root is clean. +6. **Check registered dependents** on JuliaHub and record the result in + the Phase 8 release notes draft. This is the only input to whether the + positional return-type change needs an announcement beyond the changelog. + +Acceptance gate: review the contract, deprecation list, fixture inventory, +and dependent-package findings. Run the checker against hand-enumerated +ordinary and negative examples and run the existing suite. List pending +tests with their activation phase; only the known legacy regression is an +expected failure. Record benchmark fixtures, hardware/runtime metadata to +collect, and the measurement procedure for Phase 3. + +--- + +## Phase 2: The model: `TestSpace`, rules, feasibility + +Goal: a pure, engine-independent layer that answers every question about a +space. No generation yet. Everything here is testable by hand. + +Steps: + +1. **`TestSpace`** (`space.jl`). Constructors: `TestSpace(nt::NamedTuple; + constraints)` and `TestSpace(pairs::Pair{Symbol}...; constraints)`. + Validation with messages in the user's vocabulary: names distinct; + each domain nonempty and ordered; duplicate values rejected by concrete + type and `isequal` ("parameter `tol` lists `1.0` twice"). Fields: + `names`, `values` (a tuple of vectors), `constraints`, and the tabulated + tables from step 5. Copy domains without promoting their values. + Accessors: `parameters(space)`, `arity(space)`, + `Base.length` is the full product (as `BigInt`-safe `prod`). +2. **`Invalid(x)` and `Partition(name, draw)` wrappers** are defined here + using the Phase 1 identity and collision rules. A rule or pattern sees + a partition's name (a `Symbol`). Validate domains and wrappers now; + implement negative-row rule selection in the model so feasibility and + measurement share it. Generation and realization arrive in Phase 6; + until then generation on wrapper spaces fails with an explicit + unsupported-feature error rather than applying ordinary semantics. +3. **`Constraint`** (`constraints.jl`): `scope::Tuple{Vararg{Symbol}}`, + `predicate`, `polarity` (`:forbid`/`:require`, kept for display; both + normalize to "forbidden combinations" at tabulation), `label` (reason + and/or source text). Constructors: + - `forbid(pattern::NamedTuple; reason)` — exact partial assignment. + - `forbid(names::Symbol...; reason) do values... end` and the same for + `require`. + - `forbid(f; reason)` with no names is the whole-case escape hatch: `f` + receives the complete case as a `NamedTuple`. + Construction errors name the parameter and list the space's names. + A predicate that returns anything but `Bool` is an error; an exception + thrown by a predicate is rethrown with the rule's label and arguments. +4. **`@forbid` and `@require`.** Fable's design: free identifiers that are + not in call position are parameter names; `$x` interpolates from the + caller; the source text is kept. The unknown-name error suggests + `$name` if a variable was meant. The macros produce the same + `Constraint` as step 3, so nothing downstream knows they exist. +5. **Tabulation.** For each rule, evaluate once per combination of its + scope's domains and store the forbidden index tuples in a `Set`. Above + a threshold (default 10^5 evaluations, configurable) evaluate lazily + with a memo and warn once, suggesting a narrower scope. Whole-case + rules are always lazy. Tabulation order is fixed so reports are stable. + Evaluate ordinary values (partition names included); negative-row + checks select only tables whose scopes omit the invalid parameter. +6. **Feasibility** (`feasibility.jl`). `violates(partial_idx)` fires only + when a rule's whole scope is assigned. `completable(partial_idx; limit)` + is backtracking with forward checking that returns `(true, witness)`, + `(false, nothing)`, or `:unknown` when `limit` nodes are exceeded. + Solve every constrained connected component, including components with + no assigned parameter. Cache witnesses for independent components and + combine them into a complete witness; only unconstrained parameters may + be filled freely. An unsatisfiable component makes the whole completion + infeasible. Memo keys include the assignments and active rule set, so + negative rows, deletion searches, and diagnosis cannot reuse stale + answers. Never cache an exhausted search as infeasible. + `classify(space, targets)` labels each target combination `required`, + `forbidden` (which rule), `implied` (a proven sufficient rule set found + by deletion search), or `unknown`. Claim inclusion-minimality only if + every remaining rule is verified necessary. If explanation search hits + its limit, retain the last proven sufficient set and mark minimality + unresolved; target feasibility and explanation quality are separate. +7. **`isallowed(space, case)` and `explain(space, partial)`.** Five + outcomes: `allowed`, `forbidden` (with the rule), `completable` (with a + witness), `infeasible` (with rules), `unknown`. The result prints a + sentence and carries fields. +8. **Tests** (`test/test_space.jl`, `test_constraints.jl`, + `test_feasibility.jl`): Astra's `A == B`, `B == C` example; Fable's + solver example (3 direct, 2 implied, listed by name); Opus's + os/gpu/driver example; `nothing` as a real value; the `!=` rule that + silently over-forbade in 0.4 now forbids exactly two pairs; the macro + error message; the threshold warning; `unknown` with `limit = 1`; + determinism of tabulation order. Add fixtures where the assigned + parameter is unconstrained but a separate component is unsatisfiable, + where two satisfiable components need non-default witness values, and + where a whole-case predicate connects them. Check type-preserving + identity, wrapper collisions, active rule sets, and retries after limits. + +Acceptance gate: all model tests and the full suite pass. Review executable +examples of all rule forms, the solver example's 3 direct and 2 implied +exclusions, the disconnected-component fixtures, and `unknown` followed by +a successful retry. Inspect witness validity and explanation status fields. + +--- + +## Phase 3: Engines that keep the promise + +Goal: IPOG and GND return certified designs or clear errors, terminate under +configured limits, and remove `disallow`. Random problems go green. + +Steps: + +1. **Internal request** (`request.jl`). One struct the engines consume: + `arity`, `strength`, `groups` (index tuples with their strength, base + group included), `dead(partial_idx)::Bool` (true only for proven + infeasibility, false only with a completion witness; throws a + resource-limit error on `unknown`), `must_include` as an index matrix + with `0` for unset, the feasibility budget and run-local memo, and + the classified target list from Phase 2. One entry point, + `generate(engine, request)`, returns the index matrix plus bookkeeping: + which targets were covered, how many rows came from `must_include`, + and the engine's seed. All target classifications must be resolved + before a result is returned. Validate feasibility of the whole space + even when there are no targets; a proven empty space returns an empty + design unless must-include requirements make the request impossible. +2. **IPOG.** Replace `disallow` at the three sites + (`choose_last_parameter_filter!`, `insert_tuple_into_tests_filter`, + `fill_remaining_missing_values_filter!`) with `dead`. Because every + site commits a value only when the row stays completable, the + completability invariant holds by induction; the initial combinations + and each pushed tuple are completable because the feasibility pass + removed the rest. The unconstrained `ipog` fast path stays. In + `ipog_multi_way`, copy the groups, accept ranges and tuples, and treat a + group at the base strength as a no-op instead of asserting. +3. **GND.** `allowed_argmax` uses `dead`. Progress guarantee: if no + candidate in a round covers anything, build one directly from the + first uncovered target with a stored or newly proven witness, so the + `error("Could not construct...")` path and the attempt cap go away. + `GND(; seed = 0, candidates = 50, rng = nothing)`: a fixed default + seed, `M` accepted with a deprecation warning, the seed recorded in the + result. Exhausting a witness-search budget raises the same resource-limit + error; it never drops the target or retries indefinitely. + `n_way_coverage_init` and the instrumented IPOG variant are + deleted if no test uses them. +4. **Excursions.** `build_excursion(arity, distance, base_idx, dead)` + with an explicit base; rows that are dead are dropped and *reported* + in the bookkeeping (which values never appear); a forbidden base is an + error naming the rule. +5. **Full factorial.** Enumerate candidates incrementally and retain only + accepted rows; the public return remains a materialized `TestCases`. + There is no lazy public result in this release. Before enumeration, + refuse a candidate product above `limit = 10^6`, giving its count; + report candidate and accepted counts separately. Phase 6 extends the + count to the ordinary product plus all single-invalid products. +6. **Final validation** inside `generate`: every row complete, every row + passes every applicable rule (including lazy predicates), no target + classification remains unknown, and every required target is covered. + Validate against the requested strategy: covering targets, excursion + distance, or full-factorial completeness. Phase 6 validates ordinary + and negative coverage separately. A limit raises a resource-limit error; + an invariant failure raises an internal error naming the row or target. +7. **Delete** `wrap_disallow`, the `disallow` keyword, `Counter`, + `seeds_to_integers`' sentinel handling, and documentation of `nothing` + as a partial-assignment sentinel; retain examples using it as a real + value. Single-valued parameters are allowed; a strength + larger than the parameter count is an `ArgumentError`; equal to it is + a full factorial under the applicable row policy. +8. **Turn on the random problems.** Both engines, both strengths, every + result checked by `test/checker.jl`. Close the Phase 1 issue. +9. **Performance baseline.** Benchmark 15 four-valued parameters at + strength 4 and a checked-in fixture for Fable's 12-parameter constrained + example. Record Julia version, hardware, thread count, engine seed, + compilation policy, allocations, and median of repeated warm runs; + report cold-start compilation separately. Measure the prior revision + on fixtures it handles correctly, and establish a baseline for repaired + constrained cases. Use stable CI runners to set an explicit regression + tolerance from observed variation. Ordinary correctness tests assert + completion and coverage, not machine-dependent seconds. Include limit + exhaustion and progress regressions in the deterministic suite. + +Acceptance gate: all 1,000 random problems pass for each engine at the +full multiplier, targeted regressions pass, and the full suite is green. +Review the benchmark command, environment, before/after results, and chosen +regression tolerance, plus the removal of `disallow`. Demonstrate that a +limited search raises an error rather than returning an incomplete design. + +--- + +## Phase 4: The public interface and the result type + +Goal: the 1.0 surface, with the old spellings deprecated and every +existing test rewritten against it. + +Steps: + +1. **`TestCases{T} <: AbstractVector{T}`** (`testcases.jl`). Fields: + `cases::Vector{T}`, `space`, `strategy` (`:covering`, `:excursion`, + `:full_factorial`), `strength`, `stronger`, `engine` (name and seed), + `n_must_include`, `excluded` (direct with rule labels, implied with + rules and explanation status), and `covered` (count of required targets, + from bookkeeping). Generated designs have no unknown target status; + analysis results may. Keep separate ordinary and negative bookkeeping + as specified in Phase 1. `T` is a `NamedTuple` type for named spaces and + a `Tuple` type for positional calls. Preserve domain values and their + concrete types without conversion. Homogeneous domains use concrete + field types; heterogeneous domains may use union or abstract field types. + `Any[1, 1.0]` must remain two distinct choices through generation, + collection, rule evaluation, and coverage. `collect` gives a plain vector. +2. **`covering(space; strength = 2, stronger = [], must_include = [], + engine = IPOG(), constraints = [])`** (`interface.jl`), with + `all_values`, `all_pairs`, `all_triples` as fixed strengths. Each + accepts a `TestSpace`, a `NamedTuple` of domains, `Symbol => values` + pairs, or bare positional vectors (which build a space named + `p1, p2, ...` and return tuples). `constraints =` on named domain input + builds the space; on a `TestSpace` it is an error ("constraints belong + to the space"). Expose the Phase 1 `feasibility_limit` on generation + and analysis entry points, and document resource-limit errors and retry. + Positional calls reject `constraints =`; use a named space for rules. +3. **`must_include`.** Named: `NamedTuple`s, partial allowed, an existing + `TestCases` accepted. Positional: tuples or vectors. Validated by name + and value before generation with messages that name the case, the + parameter, and the offending value; a case violating the applicable + row policy is an error. A partial case must have a proven completion; + after Phase 6, a partial case containing `Invalid` uses the negative-row + policy. Limit exhaustion is distinct from an infeasible must-include. + Must-include cases come first, in the order given; duplicates kept. + `seeds` accepted with a deprecation warning. +4. **`stronger = [(:a, :b, :c) => 3]`** validated per the contract: names + exist, distinct within a group, group strength at least the base + and at most the group size, overlapping groups combine by union, the + caller's vector untouched. A group at base strength is a no-op, matching + Phase 3. Positional: `[(1, 3, 4) => 3]`. `wayness` + accepted with a deprecation warning and translated. +5. **`excursions(space; from = nothing, distance = 1, must_include)`** + and `full_factorial(space; limit)`; `values_excursion`, + `pairs_excursion`, `triples_excursion` kept as thin aliases. The + docstring for `excursions` states its promise (within `distance` of + the base) and that it is not a covering guarantee. +6. **`show(io, ::TestCases)`.** One summary line ("5 cases · strength 2 · + IPOG · 3 parameters · 12 combinations, 5 valid" when the valid count + is cheap, otherwise just the product), then "excluded: 3 pairs + forbidden, 2 impossible under the constraints; see + report(cases)", then an aligned table truncated like a `DataFrame`. + Nothing is computed at display time that generation did not already know. +7. **Tables.** A vector of `NamedTuple`s already satisfies Tables.jl's + row-table interface; add a test that `DataFrame(cases)` and + `CSV.write` work (both are already test dependencies). +8. **Deprecations and removals.** `all_tuples`, `n_way`, `seeds`, + `wayness`, `GND(M=)` warn once via `Base.depwarn`. `generate_tuples`, + `Excursion`, `Counter` removed. `disallow` is not a keyword anywhere; + passing it hits the ordinary unknown-keyword error. +9. **Rewrite existing tests** (`test_factorial_interface.jl`, + `test_excursions.jl`, `test_full_factorial.jl`, `nonfunctional.jl`, + `cli.jl`) against the new surface. Add deterministic tests for + heterogeneous domains, partial and duplicate must-includes, overlapping + stronger groups, and preservation of caller-owned inputs. Aqua stays green. + +Acceptance gate: run generation, explanation, and display examples from +the target screen (measurement and add-on calls remain pending until their +phases). Review the transcript and deprecation warnings; confirm typed-value +preservation, Tables interoperability, targeted regressions, Aqua, and the +full suite pass. Display must perform no feasibility searches. + +--- + +## Phase 5: Measure and explain: `coverage`, `report`, `design_sizes` + +Goal: the claim is checkable in one line, and the package says what it +promised, what it left out, and what a design would cost. + +Steps: + +1. **`coverage(cases, space; strength = 2, stronger = [])`** + (`measure.jl`), for any iterable of `NamedTuple`s (or tuples for a + positional space), independent of the generator's bookkeeping. + Returns a `Coverage` with `covered`, `feasible`, `missing` (named + combinations), `unknown`, and a per-group breakdown. Separate ordinary + and negative coverage according to Phase 1; rows violating their + applicable rules contribute nothing and are listed; duplicates count + once. Negative rows never increase ordinary coverage. Implement the + result structure now and activate wrapper inputs in Phase 6. + `coverage(cases::TestCases)` reads the space and request from the + result. `coverage(cases::TestCases)` measures at `cases.strength` and + `cases.stronger`; for a result whose strength is 0 (an excursion or a + full factorial) it is an `ArgumentError` asking for `strength =`, and + `report(cases)` on such a result measures at strength + `min(2, parameter count)` and says so in its guarantee line. Prints + "covers 11 of 11 feasible pairs" or the missing list when + classification is resolved. With unknown targets, + report known counts and unresolved targets without an exact percentage + or completeness claim. +2. **`missing_interactions(cases, space; ...)`** returns + `coverage(...).missing` when classification is resolved; otherwise raise + a resource-limit error directing the caller to `coverage` for the known + missing and unresolved targets. An empty list must not hide uncertainty. +3. **`report(cases::TestCases)`**: the guarantee line, the excluded list + with attribution, bonus coverage at strength+1, the prefix curve + ("first 5 of 10 cover 73%"), and the seed. Returns a `Report` whose + fields serialize; printing is the display. This is where the + verification Astra keeps out of `show` happens. Apply the same unknown + policy to bonus coverage and prefix curves. When no strength+1 targets + exist, mark bonus coverage not applicable. +4. **`design_sizes(space; strengths = 1:3, distances = 1:2)`**: one row + per strategy with the case count, share of the valid product, and + pairs and triples covered, computed by running IPOG per strength. Full + factorial reports total and valid counts when the product is below + the Phase 3 limit, otherwise the total only. Omit shares when the valid + count is unknown; report a strategy's resource-limit status rather than + an invented case count. Skip strengths exceeding the parameter count. +5. **Top-up and extend as recipes**, tested but not new functions: + `all_pairs(space; must_include = existing)` and + `all_triples(space; must_include = cases)`. A test asserts that the + existing cases survive in order and that the result is complete. +6. **Tests**: `coverage` of hand-written cases against the checker; + `coverage` of every generated covering design is complete; `report` numbers + agree with the checker on the random problems; `design_sizes` + reproduces Fable's 10-of-81 example. Test limit exhaustion in coverage, + missing-interaction queries, bonus reports, and planning; verify that + unresolved denominators never print as exact percentages. `report` on a + one-parameter excursion and full factorial measures at strength + `min(2, parameter count)` = 1: cover this boundary in Phase 5 tests. + +Acceptance gate: review checked outputs for the solver example, a +hand-written suite with gaps, and a limited search with unresolved targets. +Run audit/top-up/extend examples and compare their results to the independent +checker. All measurement tests and the full suite pass. + +--- + +## Phase 6: `Invalid`, `Partition`, `diagnose`, `github_matrix` + +Goal: implement the four Option 4 features using the Phase 1 contract, +with matching docstrings and `diagnose` labeled experimental. + +Steps: + +1. **`Invalid(x)`** (`invalid.jl`). Generation: the covering design is + built over valid values only; then, for each invalid value `v` of + parameter `p`, build the negative targets specified in Phase 1 for the + base and all stronger groups containing `p`. Generate completions over + the other parameters' ordinary values under the applicable rules, then + insert `p = v` at its original parameter position. At strength 1, use + one witness for the empty target; do not call a strength-0 public API. + Pairwise generation pairs the invalid value with every feasible ordinary + value of every other parameter. Report infeasible targets and stop on + unresolved feasibility, exactly as for ordinary generation. The + case keeps the wrapper (`n = Invalid(-1)`) so a test body can branch + on `hasinvalid(case)`; `show` marks the rows; `report` and `coverage` + state the two guarantees separately. Rules mentioning `p` are not + applied in rows where `p` is invalid. Activate wrapper support in + generation, must-includes, full factorial, excursions, and measurement, + removing the temporary unsupported-feature errors from Phase 2. +2. **`Partition(name, draw)`** (`partition.jl`). Coverage is over names. + `realize(case; rng)` preserves tuple or named-tuple shape, substituting + `draw(rng)` for each partition; `realize(cases; rng)` maps over the vector. A + fixed-value label is `Partition(:tiny, Returns(1e-9))`. Nothing is + inferred from a function-valued domain; only the wrapper triggers this. + Preserve `Invalid` markers and ordinary values during realization; + nested wrappers remain rejected. Coverage is measured on the original + labeled cases, so callers retain those alongside realized inputs. +3. **`diagnose(cases, passed::AbstractVector{Bool}; strength)`** + (`diagnose.jl`). For sizes 1 to `strength`, list combinations present + in at least one failing case and no passing case; rank by failures + containing them, then by size, with a deterministic tie-break. Include + pass/fail counts and group candidates with identical observed occurrence + patterns as observationally indistinguishable. Every candidate is + unverified by passing cases; do not add a redundant "masked" flag. + Validate outcome length and handle no failures, all failures, and + conflicting outcomes for duplicate cases explicitly. + `followups(diagnosis)` attempts to find a valid case containing each + suspect and no other suspect using the Phase 2 witness search with + temporary forbidden tables. These isolation conditions always apply, + including when they mention an invalid parameter. Return per-suspect + status: `found` with a + witness, `inseparable` with a proven explanation, or `unknown` when the + search limit is exhausted. Nested suspects and constraints may make + isolation impossible. Prefer small changes from a failing case as a + heuristic, with no minimum-distance guarantee. Use the appropriate + ordinary or negative row policy; distinguish observational equivalence + from proven inability to isolate a suspect. Docstrings + say: experimental; hypotheses, not proof; multiple faults and + intermittent failures can confuse the ranking. +4. **`github_matrix(cases; io = stdout)`** (`export.jl`). Emits + `{"include": [...]}` using an established JSON encoder added as an + explicit dependency. Accept named rows with strings, finite real numbers, + booleans, and `nothing` (JSON null); encode Symbols as strings. + Reject unsupported values, non-finite numbers, and wrappers with a + message naming the row and field; callers explicitly map those values + to supported data first. Validate all rows before writing to `io`. + Warns above 256 jobs. The docstring shows the workflow YAML that reads + it with `fromJSON`. +5. **Tests**: check ordinary and negative guarantees with the independent + oracle, including strength 1, mixed strengths, infeasible negative + targets, wrapper collisions, and invalid-only domain rejection. + Verify `realize` determinism under a seeded `Xoshiro` and preservation + of positional shape. Opus's newton/sparse example checks ranking and + every follow-up that can be isolated; nested and constrained suspects + exercise `inseparable` and limited searches exercise `unknown`. + Parse exported JSON and compare values, including quotes, backslashes, + control characters, Unicode, nulls, empty results, and rejected + unsupported/non-finite values. Do not rely only on a fixture string. + +Acceptance gate: run all four feature examples and the complete target +screen. Review separate ordinary/negative coverage, realization output, +diagnosis statuses for found/inseparable/unknown cases, and parsed matrix +JSON. Wrapper integration tests, targeted regressions, and the full suite +pass with no wrapper feature tests pending. + +--- + +## Phase 7: Documentation and first contact + +Goal: a README that leads with the situation, a manual organized as +tutorial, how-to, explanation, and reference, and every example a doctest. + +Steps: + +1. **README first screen**: the promise sentence ("Describe the + configurations your code must handle; it tells you which combinations + your tests exercise, and supplies a compact set of additional cases + covering the rest"), Fable's decision table with Opus's three "use something else" + rows, one named example with its printed summary, install line. +2. **Tutorial** (`man/tutorial.md`): Fable's levels 0–4 in order, each + adding one concept, ending with `coverage` and `report`. +3. **How-to guides**, one page each, by job: test a function with many + options; test generic code across types (the Julia-specific pitch); + plan a CI matrix with `github_matrix`; run a simulation campaign; + audit and extend an existing suite (`coverage`, `must_include`); + diagnose a failure; test invalid inputs (a second space *and* + `Invalid`, side by side); combine with property-based testing + (`Partition`, and the no-API pattern); commit a design as data (the + lockfile pattern using `repr(collect(cases))` plus one `coverage` test). +4. **Explanation**: choosing values and oracles, moved from the paper's + "How to Use"; interaction coverage and what the evidence shows, + including the random-at-equal-budget formula; the constraint + semantics in plain words; engines page corrected (GND is shorter only + at high arity; deterministic vs seeded); IPOG page kept. +5. **Reference**: autodocs, with every exported docstring beginning + with the situation it serves ("Use when..."). Deprecated names + documented in a migration table (0.4 → 1.0), including the deleted + `disallow` and how to rewrite it as a rule. +6. **Developer pages**: the contract from Phase 1, non-goals, contributing. +7. **Agent one-pager** (`docs/src/man/agents.md`, also linked from the + README): the decision rule, three canonical patterns, and the + one-line check (`coverage`). +8. **Doctests on.** `DocMeta.setdocmeta!` and `doctest = true` in + `docs/make.jl`; a `@testitem` runs `Documenter.doctest` in the test + suite so CI fails on a stale example. + +Acceptance gate: build rendered docs (`julia --project=docs docs/make.jl`) +and run doctests and the full suite. Review the README, migration table, +and examples for limits, negative coverage, and unresolved diagnosis. +Check that no example promises a minimum case count, minimum explanation, +or unconditional follow-up isolation. + +--- + +## Phase 8: Release 0.5 + +Steps: + +1. Full matrix green (1.10, 1.13, latest release, three OSes), Aqua, + doctests, deterministic regressions, the full random-problem gate, and + benchmark results within the Phase 3 tolerance on the documented runner. +2. `CHANGELOG.md` for 0.5.0: breaking changes (return type, `disallow` + removed, `Counter` removed), deprecations with their replacements, new + API, the constraint fix with the issue number. +3. Set `version = "0.5.0"`; prepare PR `release/0.5` → `main` with the + validation results and rendered documentation. Keep it unmerged for review. +4. Draft a short Discourse announcement (the promise sentence, the + decision table, the constraint fix) for you to post or not. +5. **Pre-publication acceptance gate.** Review the changelog, release PR, + final CI results, benchmark report, and announcement draft. The user's + Phase 8 review must approve release before merge or registration. +6. **After approval:** merge the release PR, register with + `@JuliaRegistrator register` on the merge commit, and verify that TagBot + tags and docs deploy to `stable`. Report the released version and URLs; + the Discourse announcement remains a draft for you to post or not. + +Completion check: registration, tag, and stable docs correspond to the +approved release commit. Surface any external automation failure explicitly. + +--- + +## Order of work inside a phase + +Execute each phase through its acceptance gate, then present its concrete +artifacts for the user's review before beginning the next phase. Session +count is an estimate, not a completion criterion. For behavior changes, +write or update focused `@testitem`s and run them as each step lands. Run +the full suite (`julia --project -e 'using Pkg; Pkg.test()'`) at the phase +gate, and earlier when integration changes warrant it; avoid a full-suite +run after every small edit. Commit coherent changes with their checks. +Each review includes the phase PR, commands and results, example outputs, +and any explicitly pending tests assigned to a later phase. No tests remain +pending at release. If implementation conflicts with the contract, retain +the contract and raise the concrete conflict at review; do not silently +change the semantics to make a test pass. diff --git a/design/benchmark_ci_20260927.md b/design/benchmark_ci_20260927.md new file mode 100644 index 0000000..b531fd9 --- /dev/null +++ b/design/benchmark_ci_20260927.md @@ -0,0 +1,105 @@ +# UnitTestDesign benchmark + +- Date: 2026-09-27 17:57 +- Package: /home/runner/work/UnitTestDesign.jl/UnitTestDesign.jl, commit `6df2fa25343f1d48056295d43aa4760f37b87c06` +- Julia 1.13.1, x86_64-linux-gnu, Linux x86_64 +- CPU: AMD EPYC 7763 64-Core Processor, 4 logical cores, 15.6 GiB memory +- Julia threads: 1; optimization level -O2; bounds checks default +- Machine: CI runner +- Engine seed: GND `seed = 0` (a fresh `Xoshiro(0)` per call); IPOG is deterministic +- Compilation: the first call of each measurement is reported separately; timings are the median of 5 warm calls, each after `GC.gc()` + +
versioninfo() + +``` +Julia Version 1.13.1 +Commit 96ca370cf0e (2026-09-25 19:34 UTC) +Build Info: + Official https://julialang.org release +Platform Info: + OS: Linux (x86_64-linux-gnu) + CPU: 4 × AMD EPYC 7763 64-Core Processor + WORD_SIZE: 64 + LLVM: libLLVM-20.1.8 (ORCJIT, znver3) + GC: Built with stock GC +Threads: 1 default, 1 interactive, 1 GC (on 4 virtual cores) +Environment: + JULIA_PKG_SERVER_REGISTRY_PREFERENCE = eager +``` +
+ +## Fixture 1: 15 parameters × 4 values, strength 4, no rules + +`all_tuples(fill(1:4, 15)...; n_way = 4, engine)`, the same call on both revisions. The design fingerprint is `hash` of the returned rows; equal fingerprints mean equal designs. + +| Engine | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Fingerprint | +|:--|--:|--:|--:|--:|--:|--:|--:|--:| +| IPOG | 958 | 8.11 | 6.84 | 6.77–7.14 | 18025.5 | 199,402,311 | 1.42 | `f0c75cba9fa65154` | + +GND left out (`--skip-slow`). + +## Fixture 2: `bench12` (test/fixtures.jl) + +12 parameters, 331776 rows, 207360 valid, four rules. `generate(engine, Request(space; strength))` through the internal request; the public `covering` arrives in Phase 4. Queries, nodes and rule checks are the request's `feasibility.stats` for one call. Full factorial is `generate_full_factorial(Request(space; strength = 1))`. + +| Strategy | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Queries | Nodes | Rule checks | +|:--|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| +| IPOG, strength 2 | 22 | 0.434 | 0.0019 | 0.0019–0.0021 | 3.9 | 67,953 | 0.0000 | 817 | 68 | 805 | +| GND, strength 2 | 25 | 0.029 | 0.029 | 0.029–0.030 | 42.2 | 829,552 | 0.0000 | 26,837 | 68 | 35,649 | +| IPOG, strength 3 | 93 | 0.059 | 0.024 | 0.024–0.025 | 64.0 | 944,809 | 0.0000 | 6,616 | 68 | 3,697 | +| GND, strength 3 | 96 | 0.431 | 0.429 | 0.427–0.432 | 190.5 | 3,398,443 | 0.0000 | 106,528 | 68 | 134,636 | +| full factorial | 207360 | 0.120 | 0.112 | 0.111–0.113 | 191.6 | 2,644,088 | 0.0000 | 0 | 0 | 1,980,288 | + +## Fixture 3: the repaired greedy dead ends + +The four dead ends frozen in test/fixtures.jl, on which the 0.4 IPOG threw a `BoundsError`. No "before" number exists; these record the "after". + +| Fixture | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Queries | Nodes | Rule checks | +|:--|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| +| `dead_end_pairwise_1`, IPOG, strength 2 | 12 | 0.0004 | 0.0003 | 0.0002–0.0003 | 0.2 | 5,144 | 0.0000 | 71 | 96 | 559 | +| `dead_end_pairwise_1`, GND, strength 2 | 12 | 0.0016 | 0.0015 | 0.0014–0.0016 | 2.3 | 55,229 | 0.0000 | 6,642 | 165 | 1,043 | +| `dead_end_pairwise_2`, IPOG, strength 2 | 16 | 0.0005 | 0.0003 | 0.0003–0.0004 | 0.3 | 7,443 | 0.0000 | 96 | 134 | 949 | +| `dead_end_pairwise_2`, GND, strength 2 | 15 | 0.0022 | 0.0022 | 0.0021–0.0022 | 3.3 | 78,627 | 0.0000 | 9,809 | 242 | 2,035 | +| `dead_end_threeway_1`, IPOG, strength 3 | 30 | 0.0010 | 0.0008 | 0.0008–0.0008 | 1.0 | 22,876 | 0.0000 | 240 | 364 | 2,099 | +| `dead_end_threeway_1`, GND, strength 3 | 30 | 0.0067 | 0.0068 | 0.0068–0.0070 | 7.3 | 175,098 | 0.0000 | 19,667 | 752 | 4,317 | +| `dead_end_threeway_2`, IPOG, strength 3 | 60 | 0.0064 | 0.0063 | 0.0063–0.0064 | 11.8 | 218,854 | 0.0000 | 1,605 | 7,012 | 17,857 | +| `dead_end_threeway_2`, GND, strength 3 | 64 | 0.098 | 0.100 | 0.099–0.101 | 69.6 | 1,403,553 | 0.0000 | 74,884 | 48,364 | 176,720 | + +## Memory of a lazy rule's memo + +Each space gets one added whole-case rule, `forbid(case -> false)`, which is always lazy (contract §12.19), so every complete row the searches reach is memoized. The memo belongs to the request (§3.5): `memo_size(request)` counts its entries and `Base.summarysize(request.feasibility)` is the request's search state, memo included, while `Base.summarysize(space)` should not change. One space per table; the calls run in order on the same space, each with a fresh request. Times are single first calls, not medians. + +| Space | When | Cases | Time (s) | `Base.summarysize(space)` (bytes) | `memo_size(request)` (entries) | `Base.summarysize(request.feasibility)` (bytes) | Nodes | +|:--|:--|--:|--:|--:|--:|--:|--:| +| bench12 + whole-case rule | before generation | – | – | 5,440 | – | – | – | +| bench12 + whole-case rule | after IPOG, strength 2 | 22 | 0.207 | 5,440 | 982 | 978,768 | 6,800 | +| bench12 + whole-case rule | after IPOG, strength 3 | 93 | 0.054 | 5,440 | 4,870 | 5,710,288 | 54,583 | +| bench12 + whole-case rule | after GND, strength 2 | 25 | 0.180 | 5,440 | 36,558 | 25,997,320 | 146,828 | +| bench12 + whole-case rule | after GND, strength 3 | 96 | 1.07 | 5,440 | 85,380 | 96,593,480 | 497,731 | +| fixture 1 + whole-case rule | before generation | – | – | 3,864 | – | – | – | +| fixture 1 + whole-case rule | after IPOG, strength 2 | 33 | 0.250 | 3,864 | 4,451 | 3,453,768 | 23,968 | +| fixture 1 + whole-case rule | after IPOG, strength 3 | 177 | 0.618 | 3,864 | 47,666 | 54,020,808 | 359,151 | +| fixture 1 + whole-case rule | after IPOG, strength 4 | 958 | 17.7 | 3,864 | 389,417 | 392,836,232 | 3,886,299 | + +GND is not run on fixture 1 with the whole-case rule: without rules it already takes about four minutes per call at strength 4, and the rule adds a feasibility search to every value it scores. + +## Where the time goes + +Parts of the calls above, timed alone (median of 5 after one discarded call). Recorded as Phase 4/5 candidates; only the final validation changed in Phase 3 (review round 1 moved it to index space). + +| Call | Part | Time (s) | Allocated (MiB) | Note | +|:--|:--|--:|--:|:--| +| bench12 full factorial | whole call | 0.116 | 191.6 | 207360 rows of 331776 | +| bench12 full factorial | enumeration, `violates` per candidate | 0.074 | 166.2 | 64% of the call; 1,150,848 rule checks | +| bench12 full factorial | `validate_design`, `violates` per row | 0.033 | 25.3 | 29% of the call; 829,440 rule checks | +| `forbids` on a tabulated rule | one call | 26.1 ns | 32 bytes | the runtime-length `ntuple` key allocates on every check | +| bench12 full factorial | `forbids` in enumeration, estimated | 0.030 | 35.1 | 26% of the call | +| bench12 full factorial | `forbids` in validation, estimated | 0.022 | 25.3 | 19% of the call | +| bench12 full factorial | not in the call: each row to a case, `from_indices` (`to_cases`) | 0.648 | 357.5 | 560% of the call | +| bench12 full factorial | not in the call: `isallowed` on each case (the old validation) | 0.352 | 217.1 | 304% of the call | +| bench12 full factorial | ... of which `case_indices`, the value lookup | 0.285 | 109.5 | 246% of the call | +| fixture 1, IPOG | whole `generate` call | 6.88 | 18022.9 | | +| fixture 1, IPOG | classic `ipog`, the 0.4 core | 6.31 | 17278.7 | 92% | +| fixture 1, IPOG | `classify_targets`: the list of 349,440 targets | 0.772 | 504.2 | 11%; unconstrained, so nothing is excluded | +| fixture 1, IPOG | `validate_design` | 0.165 | 240.0 | 2% | + diff --git a/design/benchmark_procedure.md b/design/benchmark_procedure.md new file mode 100644 index 0000000..e1eea56 --- /dev/null +++ b/design/benchmark_procedure.md @@ -0,0 +1,117 @@ +# Benchmark fixtures and procedure for Phase 3 + +Recorded at the Phase 1 gate (2026-09-26) so that Phase 3 measures the +engines the same way before and after the constraint fix. + +## Fixtures + +1. **Unconstrained, high strength.** 15 parameters, each with 4 values, + strength 4. Both engines. This is the case where IPOG's cost is + dominated by the coverage matrix and GND's by candidate scoring. It + needs no definition beyond these numbers: no rules, and the values of + each parameter are `1:4`, so nothing is checked in for it. Its product, + 4^15, is far above the checker's `CHECK_MAX_PRODUCT` of 10^6, so its + 4-way coverage (1365 parameter groups × 256 values = 349440 targets) + must be counted directly rather than by `test/checker.jl`. +2. **Constrained, pairwise and three-way.** `bench12` in + `test/fixtures.jl`. It is a *replacement* for Fable's 12-parameter + configuration example, whose definition was not kept; only its + statistics survive (`design/interface_fable.md`, line 721). The + replacement matches every recorded statistic exactly, and + `test/test_fixtures.jl` asserts them with the checker: + + | | Recorded (Fable) | `bench12` (checker) | + |:--|--:|--:| + | Parameters | 12 | 12, arities 4,4,4,4,3,3,3,3,2,2,2,2 | + | Full factorial | 331776 | 331776 | + | Valid rows | 207360 | 207360 | + | Two-way interactions | 590 | 590 (586 feasible) | + | Uncoverable pairs | 4 | 4 | + | ... directly forbidden | 3 | 3: `(p1=2, p9=2)`, `(p5=1, p6=1)`, `(p7=1, p10=1)` | + | ... only by implication | 1, `(p1 = 2, p2 = 2)` | 1, `(p1 = 2, p2 = 2)` | + | Rules | 4, one over three parameters | 4, one over `(p1, p2, p9)` | + | Three-way targets | not recorded | 5702 feasible, 92 forbidden, 26 implied | + + The arities are the only twelve from 2 to 4 with that product, and they + give 590 pairs. Parameter `pk` takes the values `1:arity`. The fixture, + not the prose, is the reference. Both engines, strengths 2 and 3, + verified by `test/checker.jl` after each run (about 3 s pairwise and 8 s + three-way for the checker on this machine). + + **Legacy baseline** (recorded 2026-09-26, prior engines at commit + `62ab8df`, identical to `d46122d` in `src/`): Julia 1.13.1, macOS + (arm64-apple-darwin27.0.0), 8 × Apple M2, 1 thread, a laptop, not a CI + runner. + + | Engine | Strength 2 | Strength 3 | + |:--|:--|:--| + | IPOG | `BoundsError` at index 0, 2 ms warm (2.2 s cold) | `BoundsError` at index 0, 24 ms warm | + | GND (`rng = Xoshiro(0)`) | never returns: exceeded a 20,000,000-call `disallow` watchdog after 47 s | the same, after 46 s | + + IPOG builds rows with `p1 = 2, p2 = 2` and then finds no `p9` for them + (1 of its 21 rows pairwise, 6 of 91 three-way). GND loops without + raising its attempt-cap error. So the prior revision has no "before" + number for this fixture; it joins fixture 3. The IPOG failure was + asserted with `@test_throws BoundsError` in `test/test_fixtures.jl` + until Phase 3 replaced it with completeness tests for both engines; + the GND hang is recorded here only. +3. **Repaired cases.** The random problems that crash IPOG or hang GND on + the prior revision (see issue #51), `bench12`, and the four greedy + dead ends frozen in `test/fixtures.jl` (`dead_end_pairwise_1`, + `dead_end_pairwise_2`, `dead_end_threeway_1`, `dead_end_threeway_2`), + on which IPOG crashes although no target is implied. These have no + "before" number; record only the "after" so later phases can detect + regressions. + +## What to record for every measurement + +- Julia version (`VERSION`), OS, CPU model, thread count + (`Threads.nthreads()`), and whether the run is a stable CI runner or a + laptop. +- Package revision (commit hash) and the engine seed (GND default is 0). +- Compilation policy: the first call is discarded as a cold start and + reported separately; timings are the median of at least 5 warm runs + using `@timed` (or BenchmarkTools if it is added as a test dependency), + with allocations and bytes. +- The case count of the returned design, so a speedup is never bought + with a larger design unnoticed. +- Memory of lazy-rule memos (Phase 2 review round 1). Since Phase 3 + review round 1 a lazy rule's memo belongs to the operation context, the + request, and is released with it; the `TestSpace` retains nothing + (contract §3.5, §12.19). Record `Base.summarysize(space)` before and + after generation, which must not change, and, for each request, + `UnitTestDesign.memo_size(request)` and + `Base.summarysize(request.feasibility)` after generation, for two + spaces: `bench12` with an added whole-case rule, and the 15-parameter, + 4-value fixture 1 with a whole-case rule. A per-request whole-case memo + is bounded by the product of the ordinary domains + (331776 rows for `bench12`, 4^15 for fixture 1), so the second is the one + that can grow without a practical bound. The Phase 3 measurement (memo on + the space, 128 MB retained on fixture 1) is what moved the memo into the + request; a bound on the per-request memo is decided from this one. +- The search effort alongside the time: nodes and rule checks + (`Explanation.nodes`, `.evaluations`, or the `SearchStats` of the + request's feasibility searches). A separate evaluation budget (§3.3) is + added only if checks per node turn out to dominate. + +## Procedure + +1. Check out the prior revision (`main` at commit `d46122d`, before + `release/0.5`) and run fixture 1, the one it handles without crashing. + Record the table. Fixture 2's legacy baseline is the failure recorded + above. +2. Check out the Phase 3 branch and run all three fixtures. Record the + table. +3. Run the same script three times on the CI runner used for the release + and record the spread. The regression tolerance is set from that + observed variation (twice the observed spread, rounded up), not + guessed in advance. +4. Ordinary correctness tests assert completion and coverage only; no + test asserts machine-dependent seconds. Limit exhaustion and + progress-guarantee regressions go in the deterministic suite. + +The benchmark script lives at `benchmark/run.jl` and writes Markdown tables +that are pasted into the Phase 3 review. Its header gives the commands for +this revision and for the prior one. The Phase 3 results, with the 0.4 +comparison and the proposed tolerance, are in +`design/benchmark_results_phase3.md`. diff --git a/design/benchmark_results_phase3.md b/design/benchmark_results_phase3.md new file mode 100644 index 0000000..bd7e397 --- /dev/null +++ b/design/benchmark_results_phase3.md @@ -0,0 +1,266 @@ +# Phase 3 benchmark results + +Measured 2026-09-27 with `benchmark/run.jl`, following +`design/benchmark_procedure.md` (plan Phase 3 step 9). This is the first +baseline for the repaired engines; it is a laptop measurement. **A +measurement on the CI runner used for the release is still pending**, and +the regression tolerance below must be re-derived from that runner's spread +before any job enforces it. + +## How it was run + +``` +julia --project=benchmark -e 'using Pkg; Pkg.instantiate()' +julia --project=benchmark benchmark/run.jl --out run_N.md --tsv run_N.tsv # three times, N = 1, 2, 3 + +git worktree add /v04 d46122d +julia --project=/v04 benchmark/run.jl # the 0.4 code: fixture 1 only +``` + +The script needs no packages beyond the standard library: `@timed` for time, +bytes and allocation counts, `Base.summarysize` for retained memory. Each +measurement makes one first call, reported separately as the cold start, then +five warm calls, each after `GC.gc()`; the tables give the median of the warm +calls and their range. The three runs of the new revision ran back to back +in fresh processes (about 26 minutes each, most of it GND on fixture 1), then +the 0.4 run. + +## Environment + +| | | +|:--|:--| +| Package | commit `b955f50` (Phase 3 engines; `src/` unchanged by this round) | +| Prior revision | commit `d46122d` (0.4, `main` before `release/1.0`) | +| Julia | 1.13.1 (official release), `-O2`, default bounds checks, stock GC | +| OS | macOS, arm64-apple-darwin27.0.0 | +| CPU | Apple M2, 8 logical cores, 24 GiB memory | +| Threads | 1 Julia thread (1 interactive, 1 GC) | +| Machine | a laptop, not a CI runner | +| Engine seed | GND `seed = 0`: a fresh `Xoshiro(0)` per call (0.4: `GND(rng = Xoshiro(0))`, a fresh generator per call); IPOG is deterministic | +| Compilation | first call reported separately; median of 5 warm calls | + +## Fixture 1: 15 parameters × 4 values, strength 4, no rules + +The same public call on both revisions, +`all_tuples(fill(1:4, 15)...; n_way = 4, engine)`. The fingerprint is `hash` +of the returned rows. + +| Revision | Engine | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Fingerprint | +|:--|:--|--:|--:|--:|--:|--:|--:|--:|--:| +| 0.4 (`d46122d`) | IPOG | 958 | 4.89 | 3.48 | 3.43–3.49 | 17305.4 | 181,659,415 | 0.528 | `f0c75cba9fa65154` | +| Phase 3, run 1 | IPOG | 958 | 4.76 | 4.09 | 4.05–4.09 | 18055.8 | 199,432,008 | 0.725 | `f0c75cba9fa65154` | +| Phase 3, run 2 | IPOG | 958 | 4.80 | 4.04 | 3.94–4.16 | 18055.8 | 199,432,008 | 0.696 | `f0c75cba9fa65154` | +| Phase 3, run 3 | IPOG | 958 | 4.78 | 3.95 | 3.93–4.04 | 18053.1 | 199,432,008 | 0.679 | `f0c75cba9fa65154` | +| 0.4 (`d46122d`) | GND | 936 | 254.0 | 239.3 | 235.9–242.6 | 19884.3 | 76,412,022 | 0.794 | `ed6adef052056416` | +| Phase 3, run 1 | GND | 936 | 236.2 | 236.4 | 235.5–240.2 | 20883.9 | 91,331,715 | 0.982 | `ed6adef052056416` | +| Phase 3, run 2 | GND | 936 | 240.5 | 236.5 | 236.1–238.9 | 20578.1 | 91,331,715 | 0.953 | `ed6adef052056416` | +| Phase 3, run 3 | GND | 936 | 236.0 | 251.6 | 235.7–254.1 | 20640.9 | 91,331,715 | 1.09 | `ed6adef052056416` | + +Both engines return **the same design on both revisions** (equal +fingerprints and case counts: 958 rows for IPOG, 936 for GND), so the +comparison is of time for identical output. + +- **IPOG: 3.95–4.09 s after, 3.48 s before, about 14–18% slower.** The + classic unconstrained `ipog` core is unchanged; the difference is the + request's bookkeeping, which 0.4 did not have: `classify_targets` lists all + 349,440 targets (about 0.47 s) and `validate_design` checks them (about + 0.1 s), see "Where the time goes". Allocations rose 10% (182 M to 199 M). +- **GND: 236–252 s after, 239 s before, no measurable change.** The 0.4 + median lies inside the spread of the three new runs. Allocations rose 20% + (76 M to 91 M) while time did not move; the scoring over 349,440 targets + dominates both. +- The 0.4 numbers are one run, so their own spread is unknown; the new + revision's spread is in the last section. On this machine the IPOG + difference is larger than the IPOG spread (3.3%) but within the proposed + 35% tolerance. + +## Fixture 2: `bench12` + +`generate(engine, Request(space; strength))` on `test_space(bench12)` +(test/fixtures.jl), through the internal request; the public `covering` +arrives in Phase 4. Queries, nodes and rule checks are the request's +`feasibility.stats` after one call. Full factorial is +`generate_full_factorial(Request(space; strength = 1))`. Run 1 of 3: + +| Strategy | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Queries | Nodes | Rule checks | +|:--|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| +| IPOG, strength 2 | 22 | 0.248 | 0.0014 | 0.0014–0.0014 | 4.0 | 69,176 | 0.0000 | 817 | 68 | 717 | +| GND, strength 2 | 25 | 0.021 | 0.021 | 0.021–0.022 | 42.4 | 830,959 | 0.0000 | 26,837 | 68 | 35,549 | +| IPOG, strength 3 | 93 | 0.036 | 0.017 | 0.017–0.018 | 64.4 | 949,596 | 0.0000 | 6,616 | 68 | 3,325 | +| GND, strength 3 | 96 | 0.296 | 0.296 | 0.296–0.297 | 191.3 | 3,403,432 | 0.0000 | 106,528 | 68 | 134,252 | +| full factorial | 207360 | 0.820 | 0.831 | 0.826–0.834 | 742.1 | 14,589,171 | 0.056 | 0 | 0 | 1,150,848 | + +The first IPOG call (0.25 s) carries the compilation of the request and +feasibility code; later first calls are nearly warm. + +These agree with what the two engine agents reported during Phase 3: IPOG +22 rows in 1.4 ms pairwise and 93 rows in 22.6 ms three-way (here 16–18 ms); +GND 25 rows in 20 ms and 96 rows in 0.32 s (here 20–22 ms and 0.29–0.31 s); +full factorial 0.67 s, mostly validation (here 0.74–0.87 s across the three +runs, 92–93% of it in `validate_design`). Every design is checked complete by +`test/checker.jl` in the test suite (test_parameter_order.jl, +test_greedy_tuples.jl, test_fixtures.jl); the benchmark times only. + +**Before:** no baseline. The 0.4 IPOG throws a `BoundsError` at both +strengths and the 0.4 GND does not return (the legacy baseline recorded in +`design/benchmark_procedure.md`, fixture 2). + +## Fixture 3: the repaired greedy dead ends + +The four dead ends frozen in test/fixtures.jl, on which the 0.4 IPOG threw a +`BoundsError`. There is no "before"; these are the "after" for later phases. +Run 1 of 3: + +| Fixture | Cases | First call (s) | Warm median (s) | Warm min–max (s) | Allocated (MiB) | Allocations | GC (s) | Queries | Nodes | Rule checks | +|:--|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| +| `dead_end_pairwise_1`, IPOG, strength 2 | 12 | 0.060 | 0.0002 | 0.0002–0.0002 | 0.2 | 5,611 | 0.0000 | 71 | 96 | 511 | +| `dead_end_pairwise_1`, GND, strength 2 | 12 | 0.0011 | 0.0011 | 0.0011–0.0011 | 2.3 | 55,696 | 0.0000 | 6,642 | 165 | 995 | +| `dead_end_pairwise_2`, IPOG, strength 2 | 16 | 0.0003 | 0.0002 | 0.0002–0.0002 | 0.3 | 8,085 | 0.0000 | 96 | 134 | 885 | +| `dead_end_pairwise_2`, GND, strength 2 | 15 | 0.0016 | 0.0016 | 0.0016–0.0016 | 3.4 | 79,228 | 0.0000 | 9,809 | 242 | 1,975 | +| `dead_end_threeway_1`, IPOG, strength 3 | 30 | 0.060 | 0.0006 | 0.0006–0.0006 | 1.1 | 24,101 | 0.0000 | 240 | 364 | 1,979 | +| `dead_end_threeway_1`, GND, strength 3 | 30 | 0.0051 | 0.0051 | 0.0050–0.0051 | 7.3 | 176,324 | 0.0000 | 19,667 | 752 | 4,197 | +| `dead_end_threeway_2`, IPOG, strength 3 | 60 | 0.063 | 0.0043 | 0.0042–0.0043 | 12.0 | 221,701 | 0.0000 | 1,605 | 7,012 | 17,617 | +| `dead_end_threeway_2`, GND, strength 3 | 64 | 0.068 | 0.068 | 0.068–0.068 | 69.9 | 1,406,590 | 0.0000 | 74,884 | 48,364 | 176,464 | + +The random problems of the gate are the rest of fixture 3. At multiplier 1.0 +(500 problems per strength, the fixed seeds), generation alone takes, summed +over the 500 problems: pairwise, IPOG 0.33 s and GND 1.8 s; three-way, IPOG +0.92 s and GND 11.5 s. The checker takes about 2.4–3.9 s per engine and +strength on top, and drawing the problems (which runs the checker once more) +about 3.7–4.0 s. The two gate test items take 54 s together in the test +runner. + +## Retained memory of a lazy rule's memo + +> This section records the Phase 3 state. In Phase 3 review round 1 the memo +> moved from the `TestSpace` into the request (contract §12.19); see [Memory +> of a lazy rule's memo](benchmark_ci_20260927.md#memory-of-a-lazy-rules-memo) +> for the measurement under the current design. + +Contract §12.19 keeps a lazy rule's memo on the `TestSpace`, so it outlives +every call. Each space below gets one added whole-case rule, +`forbid(case -> false)`, which is always lazy; the calls run in order on the +same space, so the memo accumulates. Times are single first calls. Identical +in all three runs except the times. + +| Space | When | Cases | Time (s) | `Base.summarysize(space)` (bytes) | `memo_size(space)` (entries) | Nodes | +|:--|:--|--:|--:|--:|--:|--:| +| bench12 + whole-case rule | before generation | – | – | 5,696 | 0 | – | +| bench12 + whole-case rule | after IPOG, strength 2 | 22 | 0.131 | 407,104 | 982 | 6,800 | +| bench12 + whole-case rule | after IPOG, strength 3 | 93 | 0.039 | 1,611,328 | 5,036 | 54,583 | +| bench12 + whole-case rule | after GND, strength 2 | 25 | 0.148 | 6,428,224 | 37,714 | 146,828 | +| bench12 + whole-case rule | after GND, strength 3 | 96 | 0.730 | 25,695,808 | 97,813 | 497,731 | +| fixture 1 + whole-case rule | before generation | – | – | 4,144 | 0 | – | +| fixture 1 + whole-case rule | after IPOG, strength 2 | 33 | 0.214 | 2,002,992 | 4,451 | 23,968 | +| fixture 1 + whole-case rule | after IPOG, strength 3 | 177 | 0.380 | 31,985,712 | 48,517 | 359,151 | +| fixture 1 + whole-case rule | after IPOG, strength 4 | 958 | 9.43 | 127,930,416 | 394,410 | 3,886,299 | + +GND was not run on fixture 1 with the whole-case rule: unconstrained it +already takes four minutes per call at strength 4, and the rule adds a +feasibility search for every value it scores. + +Reading: one memo entry costs about 260–330 bytes here (a 12- or 15-integer +tuple key and a `Bool` in a `Dict`). On `bench12` the memo reaches 98,000 of +its 331,776 possible entries (26 MB) after four calls. On fixture 1 one +strength-4 IPOG call leaves 394,000 entries and 128 MB on the space, against +a bound of 4^15 ≈ 1.07 × 10^9 entries (hundreds of GB). The memo therefore +grows with the work done, not with the product, and a long-lived space that +serves several high-strength calls can hold hundreds of MB. This is the +measurement the procedure asked for before deciding whether to keep the memo +on the space, bound it, or move it into the request; that decision is left to +the Phase 3 review. + +## Where the time goes + +Parts of the calls above, timed alone (median of 5 after one discarded call; +the parts are timed separately, so they need not sum to the whole). Run 1 of +3; runs 2 and 3 give the same shares within a few points. + +| Call | Part | Time (s) | Allocated (MiB) | Note | +|:--|:--|--:|--:|:--| +| bench12 full factorial | whole call | 0.797 | 736.8 | 207360 rows of 331776 | +| bench12 full factorial | enumeration, `violates` per candidate | 0.053 | 168.5 | 7% of the call; 1,150,848 rule checks | +| bench12 full factorial | `validate_design`, `isallowed` per row | 0.740 | 568.3 | 93% of the call; 829,440 rule checks | +| `forbids` on a tabulated rule | one call | 19.9 ns | 32 bytes | the runtime-length `ntuple` key allocates on every check | +| bench12 full factorial | `forbids` in enumeration, estimated | 0.023 | 35.1 | 3% of the call | +| bench12 full factorial | `forbids` in validation, estimated | 0.017 | 25.3 | 2% of the call | +| bench12 full factorial | validation: each row to a case, `from_indices` | 0.413 | 357.5 | 52% of the call | +| bench12 full factorial | validation: `isallowed` on each case | 0.296 | 217.1 | 37% of the call | +| bench12 full factorial | ... of which `case_indices`, the value lookup | 0.257 | 109.5 | 32% of the call | +| fixture 1, IPOG | whole `generate` call | 4.21 | 18059.0 | | +| fixture 1, IPOG | classic `ipog`, the 0.4 core | 3.70 | 17310.6 | 88% | +| fixture 1, IPOG | `classify_targets`: the list of 349,440 targets | 0.476 | 506.7 | 11%; unconstrained, so nothing is excluded | +| fixture 1, IPOG | `validate_design` | 0.096 | 242.1 | 2% | + +## Run-to-run spread and the proposed tolerance + +Warm medians of the three runs of the new revision: + +| Measurement | Cases | Run 1 (s) | Run 2 (s) | Run 3 (s) | Spread | Widest warm min–max within a run | +|:--|--:|--:|--:|--:|--:|--:| +| fixture 1, IPOG, strength 4 | 958 | 4.09 | 4.04 | 3.95 | 3.3% | 5.6% | +| fixture 1, GND, strength 4 | 936 | 236.4 | 236.5 | 251.6 | 6.4% | 7.8% | +| bench12, IPOG, strength 2 | 22 | 0.0014 | 0.0013 | 0.0014 | 10.4% | 60.7% | +| bench12, GND, strength 2 | 25 | 0.021 | 0.020 | 0.022 | 13.0% | 6.7% | +| bench12, IPOG, strength 3 | 93 | 0.017 | 0.016 | 0.018 | 8.6% | 7.4% | +| bench12, GND, strength 3 | 96 | 0.296 | 0.290 | 0.310 | 6.9% | 1.2% | +| bench12, full factorial | 207360 | 0.831 | 0.744 | 0.865 | 16.2% | 1.9% | +| dead_end_pairwise_1, IPOG, strength 2 | 12 | 0.0002 | 0.0002 | 0.0002 | 44.9% | 26.3% | +| dead_end_pairwise_1, GND, strength 2 | 12 | 0.0011 | 0.0010 | 0.0011 | 12.1% | 6.5% | +| dead_end_pairwise_2, IPOG, strength 2 | 16 | 0.0002 | 0.0002 | 0.0003 | 29.7% | 32.5% | +| dead_end_pairwise_2, GND, strength 2 | 15 | 0.0016 | 0.0014 | 0.0015 | 9.9% | 11.1% | +| dead_end_threeway_1, IPOG, strength 3 | 30 | 0.0006 | 0.0005 | 0.0006 | 11.6% | 21.1% | +| dead_end_threeway_1, GND, strength 3 | 30 | 0.0051 | 0.0047 | 0.0050 | 8.7% | 17.5% | +| dead_end_threeway_2, IPOG, strength 3 | 60 | 0.0043 | 0.0040 | 0.0045 | 11.8% | 19.7% | +| dead_end_threeway_2, GND, strength 3 | 64 | 0.068 | 0.064 | 0.071 | 9.9% | 2.1% | + +Spread is (largest − smallest) / smallest of the three medians. Case counts, +allocation counts, rule checks, nodes and fingerprints were identical in all +three runs; allocated bytes varied by under 2%. + +Proposed regression tolerance, **for this laptop only**: + +- **Time, calls of 10 ms or more: 35%.** The largest spread among them is + 16.2% (bench12 full factorial; GND on fixture 1 is 6.4%), and twice that, + rounded up to the next 5%, is 35%. A later run is a regression when its + warm median exceeds the Phase 3 median by more than 35%. +- **Time, calls under 10 ms: not gated.** Their spread reaches 45% (the + sub-millisecond IPOG dead ends), so a tolerance of twice that would catch + nothing a reviewer would not see by eye. +- **Allocation counts: exact, per Julia version.** They were identical in + every run, so any change is a code change and should be explained in the + review rather than tolerated. + +These numbers come from one laptop on one afternoon. The procedure sets the +tolerance from three runs on the stable CI runner used for the release; that +measurement has not been made, and until it is, the tolerance above is a +proposal, not a gate. + +## Phase 4/5 candidates (measured, not optimized) + +1. **`validate_design` round-trips every row through values.** On `bench12` + full factorial it is 92–93% of the call (0.67–0.77 s of 0.73–0.83 s, + timed alone). The engine agents attributed this to per-row `isallowed`; the breakdown + shows the rule checks themselves are small: building a `NamedTuple` per + row (`from_indices`) is about 52–57% of the call and reading it back + (`case_indices`, a lookup of each value by identity) about 27–32%. + Checking the index row against the tables directly (the enumeration's + `violates` does the same work in 0.05 s) would remove most of it. + Whether final validation should stay in the caller's vocabulary as an + independent check is a design question for the review. +2. **`forbids` allocates a runtime-length `ntuple` key on every check**: + about 17–20 ns and 32 bytes per call. It is only 2–3% of the full + factorial each in enumeration and validation, so it matters less than + item 1, but it is on every feasibility search's path (the rule checks + column above: 1.15 million in one full factorial, 134,000 in GND + three-way on `bench12`). +3. **Unconstrained requests still build the target list.** On fixture 1 the + new IPOG is about 0.5–0.6 s slower than 0.4 in the same call; the + `classify_targets` list of 349,440 targets (0.47 s, 506 MiB) and + `validate_design` (0.1 s) account for it. The design itself is identical + (same fingerprint). Counting targets instead of listing them when no rule + excludes anything, or validating coverage with the engine's coverage + matrix, would close most of the gap. +4. **GND on fixture 1 costs four minutes per call**, the same as 0.4: the + candidate scoring over 349,440 targets dominates, as the procedure + anticipated. Nothing in Phase 3 changed it. diff --git a/design/interface_astra.md b/design/interface_astra.md new file mode 100644 index 0000000..ff435b1 --- /dev/null +++ b/design/interface_astra.md @@ -0,0 +1,471 @@ +# A proposed interface for UnitTestDesign.jl + +This is Astra's best guess at an interface, intended to be compared with other +proposals. It describes proposed behavior, not the current implementation. The +examples are interface sketches; they are not executable against today's package. + +## Purpose and scope + +The package should help someone answer: + +> I have several choices to test together. Which cases will exercise their +> interactions without running every possible configuration? + +The strongest application is a test whose configurations are meaningful and whose +execution is expensive enough to justify deliberate selection. Exhaustive loops +remain a good choice for small, cheap spaces. Randomized and property-based tests +remain useful for exploring concrete inputs. This interface should fit inside +those approaches rather than require adopting a new testing framework. + +The proposal does not assume that a larger feature set will create demand. Its +central bet is that named choices, understandable constraints, and inspectable +coverage make the existing algorithms useful with little additional work. + +The author supplies three things: + +1. Finite choices, with names that mean something in the problem domain. +2. Rules describing which choices can occur together, when needed. +3. Ordinary Julia code to construct inputs and check behavior. + +The library chooses configurations. It does not infer correct behavior, provide +a test runner, or require users to describe their assertions in another language. + +## 1. The small example should be the everyday interface + +```julia +using UnitTestDesign +using Test +using Random + +choices = ( + T = (Float32, Float64), + storage = (:dense, :view), + n = (0, 1, 17), + pattern = (:zeros, :alternating, :random), +) + +cases = all_pairs(choices) + +@testset "my_sum $case" for case in cases + rng = Xoshiro(1234) + A = make_input(case; rng) # A fresh input for this configuration. + @test my_sum(A) ≈ reference_sum(A) +end +``` + +Here `make_input`, `my_sum`, and `reference_sum` belong to the user's tests. The +example assumes an appropriate reference and comparison for those inputs. + +Each case is a `NamedTuple` with the same names and order as `choices`: + +```julia +(T = Float32, storage = :view, n = 0, pattern = :zeros) +``` + +Names describe test dimensions, which need not correspond one-to-one with +function arguments. A case can describe how to construct an array, an entire +simulation, or a file. An author can pass `case...` as keywords when those names +do match a function, but that is not required. + +Keep the familiar generation functions: + +```julia +all_values(choices) +all_pairs(choices) +all_triples(choices) +all_tuples(choices; n_way = 4) +full_factorial(choices) +``` + +`all_pairs` means that every feasible pair of factor values appears together in +at least one returned case. It does not mean every pair of complete test cases, +balanced frequencies, a minimum-size suite, or coverage of every program path. + +## 2. Introduce a model object only when there is something to reuse + +Simple calls accept a named tuple directly. `TestSpace` packages the same choices +with constraints for repeated generation and inspection: + +```julia +space = TestSpace(choices) +cases = all_pairs(space) +``` + +`all_pairs(choices)` is shorthand for `all_pairs(TestSpace(choices))`. + +`TestSpace` contains choices and validity rules. Coverage strength, mandatory +cases, and algorithm selection belong to the generation request. The same space +can therefore support pairwise CI tests and a stronger release suite. + +Domains are finite, nonempty, ordered collections. Singleton domains are allowed; +users should not have to remove fixed settings from a model. `all_values` supports +a one-factor model. A requested strength greater than the number of factors is a +clear validation error. Empty models and empty domains are errors in this first +interface. + +Values can be Julia types, symbols, numbers, tuples, or other Julia objects. +Selecting a value must preserve its type. Domain lookup and exact-pattern matching +require both the same concrete type and `isequal` values. Thus `Int8(1)`, `Int64(1)`, +and `1.0` can be distinct test choices. Duplicate choices under that relation are +rejected with a useful explanation. Internally, generation works with choice +indices rather than requiring values to be sortable. + +The model copies the domain containers and does not mutate caller configuration. +Selected mutable values are not automatically deep-copied. For tests that mutate +inputs, choose descriptions or factories and construct fresh fixtures per case. +Mutating a domain value after constructing the model is outside the contract. + +## 3. Constraints describe complete cases + +The user-facing rule is: + +> Describe which complete configurations are allowed. You never need to handle +> a placeholder for an input that the generator has not chosen yet. + +There are two primary forms: forbidden patterns and required relationships. + +### Forbidden patterns + +```julia +space = TestSpace( + ( + format = (:text, :binary), + compression = (:none, :gzip), + mode = (:buffered, :streaming), + seekable = (false, true), + ); + rules = [ + forbid( + (format = :text, compression = :gzip); + reason = "This text format does not support gzip", + ), + forbid( + (mode = :streaming, seekable = true); + reason = "Streaming mode does not support seeking", + ), + ], +) +``` + +Within one pattern, all listed choices must match for it to forbid a case. Factors +omitted from the pattern can take any value. Matching any forbidden pattern is +enough to reject a case. + +| Format | Compression | First rule | +| --- | --- | --- | +| `:text` | `:none` | Allows | +| `:text` | `:gzip` | Forbids | +| `:binary` | `:none` | Allows | +| `:binary` | `:gzip` | Allows | + +Patterns use exact values. A tuple-valued choice is still one value, not shorthand +for alternatives. More complicated conditions use a predicate. The pattern is a +positional named tuple so that metadata such as `reason` cannot collide with a +factor named `reason`. + +### Required relationships + +```julia +square = require((:rows, :columns); reason = "Matrix must be square") do rows, columns + rows == columns +end + +space = TestSpace( + (rows = (0, 1, 8), columns = (0, 1, 8), T = (Float32, Float64)); + rules = [square], +) +``` + +The callback receives concrete values in the order of the named factors. All +required relationships must return `true` for an allowed case. Julia's do-block +syntax puts the predicate first in the underlying function call: + +```julia +require(predicate, (:rows, :columns); reason = "Matrix must be square") +``` + +Predicates must be deterministic and return `Bool`. They can be evaluated more +than once, in an unspecified order. They should not mutate values or depend on +changing external state. A thrown exception is reported as a rule-evaluation +error, with its original cause and argument values; it does not silently mean +"forbidden." + +An escape hatch accepts an ordinary complete-case predicate: + +```julia +rule = require(; reason = "Configuration fits the memory budget") do case + estimated_bytes(case) <= memory_budget +end +``` + +The estimate and budget are user code. This form receives a complete named tuple. +Small scopes are preferable when natural because they permit earlier pruning and +more useful explanations. They are not a prerequisite for correctness. + +There is no public partial-value sentinel. `nothing` and `missing` may themselves +be real domain values; their interpretation belongs to the user's rule. A +predicate returning `missing` instead of `Bool` is an error. + +### What constraints mean for coverage + +A combination requires coverage if and only if it occurs in at least one complete +case satisfying every rule. + +For example, suppose `A`, `B`, and `C` each have choices `(1, 2)`, with rules +`A == B` and `B == C`. The only complete cases are `(1, 1, 1)` and `(2, 2, 2)`. +The pair `A = 1, C = 2` is infeasible and does not need coverage. The user does not +need to supply the implied relationship `A == C`. + +The implementation must reason about valid completions. Merely finding no +immediate violation in a partial case is not sufficient. Likewise, generating +unconstrained cases and deleting forbidden rows cannot establish constrained +coverage: removing a row can erase the only occurrence of another feasible pair. + +This semantic promise is stronger than the current constraint implementation. +It is an implementation requirement, not something a nicer callback fixes by +itself. Backtracking over finite domains is a possible starting point; a solver +could support larger models or a restricted declarative rule language later. +Arbitrary Julia predicates are not promised automatic translation into a solver. + +If a search limit prevents establishing feasibility or completing a suite, the +operation reports that limit. It must not label an unresolved combination +infeasible or return a successful result claiming full coverage. + +### Invalid inputs can still be valuable tests + +Rules define the population for this particular test. They do not describe every +input the application might receive. + +For a numerical-correctness test, exclude configurations the operation does not +support. For a rejection test, deliberately generate those configurations and +assert the documented exception. The documentation should show both uses so that +"forbidden" does not accidentally become "never test this error path." + +## 4. Let users inspect rules without generating a suite + +```julia +isallowed(space, complete_case) +explain(space, complete_case) +explain(space, (format = :text, compression = :gzip)) +``` + +`isallowed` accepts only complete cases. `explain` also accepts partial assignments +and checks whether a valid completion exists. It returns a structured explanation +whose display is readable in the REPL. + +Illustrative output: + +```text +No valid completion for: + format = :text + compression = :gzip + +Conflicts with: + This text format does not support gzip +``` + +The result distinguishes `:allowed`, `:forbidden`, `:completable`, +`:infeasible`, and `:unknown`. A completable partial assignment includes one +witness completion. `:unknown` means a search limit prevented a conclusion. + +A directly violated rule can be named immediately. An explanation involving +several interacting rules may be less specific; the first implementation need +not find a smallest conflicting subset. It must not invent a causal explanation. +Unknown names and out-of-domain values are input errors, not constraint failures. + +Model construction checks names, domains, and rule structure. It does not imply a +potentially expensive proof that the whole model is satisfiable. Generating from +an unsatisfiable model produces a model-level error, not an empty successful suite +or an indexing error. + +## 5. Make generated cases ordinary to consume and coverage explicit to inspect + +Named generation returns a `TestCases` collection supporting `length`, iteration, +and indexing. Its elements are named tuples. The collection retains its model +and generation request, including algorithm information, so users can inspect +what it was intended to cover. It does not track test execution. + +```julia +cases = all_pairs(space) +cases[1] +collect(cases) # Ordinary vector of named tuples. +report = coverage(cases) +``` + +`coverage(cases)` independently measures the returned rows against the stored +request. For cases from other sources, the request is explicit: + +```julia +report = coverage(space, existing_cases; n_way = 2) +``` + +A report includes: + +- Requested interaction strength and any stronger groups. +- Number of rows examined and any invalid rows. +- Feasible interactions required, covered, and uncovered, when established. +- A bounded sample of uncovered interactions, expressed using factor names. +- Whether checking completed, and whether full coverage was established. + +An invalid row contributes no coverage. A valid model with incomplete rows is not +reported complete merely because those rows happened to cover some projections. +If checking reaches a resource limit, the report exposes unresolved work and does +not manufacture an exact percentage. Explicit coverage verification can itself +be expensive; printing a collection must not silently perform it. + +This operation measures the supplied configurations, not passed assertions, code +coverage, or proven correctness. A user who skips cases during execution must +pass the cases actually exercised to measure executed configuration coverage. + +The small default display should show row count, factor names, and requested +strength. It should not print a wall of internal matrices or unverified metrics. + +## 6. Build on tests the user already values + +```julia +regressions = [ + (T = Float32, storage = :view, n = 0, pattern = :zeros), +] + +cases = all_pairs(choices; seeds = regressions) +``` + +Seeds are complete mandatory cases. They are checked against the domains and +rules before generation and retained in supplied order, followed by additional +cases. Invalid seeds produce an explanation; they are never silently discarded. +Duplicate seed configurations are retained, preserving the user's requested +list, but count only once toward distinct interaction coverage. Repetition for +statistical testing is ordinarily better expressed in the execution loop. + +This supports a useful workflow without a separate suite-management subsystem: + +1. Record the configurations of existing tests. +2. Inspect their coverage. +3. Supply them as seeds to generate a completed suite. + +Completion here means achieving the requested interaction coverage. It does not +promise the mathematically smallest possible number of additions. + +## 7. Expose stronger coverage without positional bookkeeping + +```julia +cases = all_pairs( + choices; + extra = [strength(3, (:T, :storage, :n))], +) +``` + +This requests pairwise coverage globally and three-way coverage within the named +group. `strength(3, (:a, :b, :c, :d))` would cover every feasible triple within +that four-factor group. Group requirements are combined by union; overlapping +requirements do not multiply obligations. + +Names must exist, names within a group must be distinct, and a group's strength +cannot exceed its size. The base and extra requirements must be represented in +the coverage report. Generating cases must not modify the supplied groups. + +Keep algorithm selection optional and retain the existing engine types: + +```julia +cases = all_pairs(space; engine = IPOG()) +cases = all_pairs(space; engine = GND(rng = Xoshiro(1234))) +``` + +The default should be deterministic for a fixed model and implementation version. +Exact ordering across future versions is not a promise. Explicitly saved cases +are stronger regression records than a generator seed alone. + +An engine that cannot satisfy the requested semantics must report that limitation +or use a documented correct fallback. Choosing a different engine must not weaken +the meaning of "all feasible pairs." + +## 8. Keep excursions visibly distinct + +```julia +cases = pairs_excursion(space) +``` + +This explores configurations differing from the baseline in at most two factors. +The baseline is the first value in each domain, matching the existing convention. +Filtering those excursions by the model's rules does not promise global pairwise +coverage. An invalid baseline should be reported clearly. + +For named models, use the explicit excursion functions. Do not encourage +`all_pairs(space; engine = Excursion())`, because that spelling suggests the same +coverage contract as the covering generators. Existing positional behavior can +remain for compatibility while its different semantics are documented. + +## 9. Fit randomized and property-based tests through ordinary composition + +A case can choose a category such as `:near_zero` or `:ill_conditioned`. User code +generates concrete values within that category and runs a property or reference +comparison. More repetitions explore additional examples under the same +configuration coverage. + +This first interface needs no property-testing dependency or custom assertion +macro. It also makes no promise to shrink concrete failures. If a property-testing +integration proves useful, it should reuse the named configurations and coverage +semantics rather than introduce a second model language. + +Human users and AI assistants receive the same interface: normal Julia data, +structured errors, named uncovered interactions, and ordinary test code. No +AI-specific API is required. Factor selection and assertions remain explicit +assumptions that can be reviewed. + +## 10. Compatibility and a deliberately limited first implementation + +Retain existing positional calls such as: + +```julia +all_pairs([1, 2, 3], ["low", "high"], [false, true]) +``` + +They retain their existing return convention. Dispatch on `NamedTuple` or +`TestSpace` selects the new named interface. The old `disallow` callback is a +legacy interface with partial-input semantics; it must not silently acquire new +meaning. The proposed `rules` interface is separate and explicitly documented. +Similarly, named groups replace positional `wayness` only in the new interface. + +The initial surface is: + +| Operation | Purpose | +| --- | --- | +| `TestSpace` | Reusable finite choices and validity rules | +| `forbid`, `require` | Exact forbidden patterns and ordinary Julia relationships | +| Existing generation names | Select cases at a requested strength | +| `strength` | Describe an additional coverage group by name | +| `isallowed`, `explain` | Inspect validity and possible completions | +| `coverage` | Measure coverage of generated or existing configurations | + +The first implementation should omit a custom runner, macro language, automatic +model inference, CLI, general serialization framework, distributed scheduling, +automatic shrinking, and a large catalog of domain-specific fixtures. These may +be useful later, but none is required to test the central interface hypothesis. + +## Questions for competing proposals + +These are the decisions on which an alternative could reasonably do better: + +1. Is a named overload of `all_pairs` easier to discover than a new entry point + such as `covering_cases`? The proposal favors familiarity over a fresh naming + scheme. +2. Are `forbid` patterns and scoped `require` predicates worth two concepts, or + would one complete-case predicate be enough? The proposal favors readable + common exclusions while keeping arbitrary Julia available. +3. Is `TestCases` with retained metadata worth more complexity than returning a + plain vector? The proposal favors inspectability, with `collect` as the exit. +4. Can correct constrained generation be acceptably fast with ordinary Julia + predicates? If not, should the package limit supported constraints, introduce + a declarative subset, or make the cost of opaque predicates more visible? +5. Is coverage inspection and completion useful enough to justify building a + model? It still requires the user to identify factors; the proposal does not + remove that cost. +6. Does this interface improve a real test suite enough that its author would + keep it after comparing exhaustive loops and random selection under the same + budget? + +The strongest first demonstration would use an existing expensive Julia test: +express its choices and rules, measure existing coverage, complete the missing +interactions, and run the existing assertions. The evaluation should include +authoring effort, execution cost, and fault detection. A smaller row count alone +would not establish that the interface is worth adopting. diff --git a/design/interface_astra.pdf b/design/interface_astra.pdf new file mode 100644 index 0000000..90fd053 Binary files /dev/null and b/design/interface_astra.pdf differ diff --git a/design/interface_fable.md b/design/interface_fable.md new file mode 100644 index 0000000..ce60e51 --- /dev/null +++ b/design/interface_fable.md @@ -0,0 +1,728 @@ +# UnitTestDesign.jl: a proposal for the next interface + +Prepared 2026-09-18 against branch `feature/upgrade` (package version 0.4.0). +Everything marked *measured* was run against the current code with Julia 1.12.5; +the probe scripts are summarized in Appendix A and the prototype in Appendix B. + +## 1. Summary + +The package does one valuable thing well: given a few representative values +per argument, it picks a small set of argument combinations that still +exercises every pair (or triple) of values. The simple call is fine. Everything +past the simple call is where users fall off, and the first thing they fall +over is constraints. This proposal makes three changes and a handful of +smaller ones. + +1. **Constraints become declarations over named parameters, evaluated only on + complete values.** `@forbid mode == :fast && solver != :none` replaces + `disallow = (m, s, t) -> ...`. The library evaluates each constraint once + over the product of the parameters it mentions, turns it into a table of + forbidden combinations, and never calls user code during search. Interactions + that the constraints make impossible are detected, dropped from the coverage + goal, and reported instead of crashing IPOG or hanging GND (both measured + today). + +2. **Results become typed tuples or named tuples inside a `TestCases` vector + that prints its own summary and is a table.** `f(case...)`, `f(; case...)`, + `(; mode, solver) = case`, `DataFrame(cases)`, and `@testset "$case" for case + in cases` all work without glue code, and the REPL shows "10 cases, pairwise, + of 81 combinations" so the tradeoff is visible. + +3. **The entry point leads with the decision, not the algorithm.** The README + and docs answer "when do I reach for this instead of property-based testing, + random testing, or `Iterators.product`?" in the first screen, and two new + functions back the answer with numbers: `design_sizes(space)` shows what each + strategy would cost, and `missing_interactions(existing_cases, space)` tells + a user what their hand-written tests already miss, so they can *extend* a + suite instead of replacing it. + +Smaller changes: `strength` and `stronger` replace `n_way` and the index-based +`wayness` dictionary; seeds may be partial; excursions take an explicit base +case; single-valued parameters are allowed; `Counter`, `generate_tuples`, and +the `Excursion` engine leave the public surface. + +## 2. What exists today + +### 2.1 Public surface + +| Exported | Role | +|---|---| +| `all_values`, `all_pairs`, `all_triples`, `all_tuples(...; n_way)` | covering designs at strength 1, 2, 3, or t | +| `values_excursion`, `pairs_excursion`, `triples_excursion` | one-, two-, three-parameter walks away from the first value of every parameter | +| `full_factorial` | every combination, optionally filtered | +| `IPOG`, `GND`, `Excursion` | engine tags passed as `engine =` | +| `generate_tuples` | the internal dispatch point, exported by accident | + +Every design function shares keywords `engine`, `disallow`, `seeds`, `wayness`, +`Counter` (`src/factorial_interface.jl:209`). + +### 2.2 How a call flows + +`all_tuples` validates the arguments and calls `generate_tuples(engine, ...)`. +That function translates values into 1-based integer indices, so the engines +only ever see an `arity` vector such as `[3, 3, 3, 3]`. Three translations +happen at that boundary: + +- `wrap_disallow` (`factorial_interface.jl:81`) wraps the user's predicate so an + integer vector becomes values, mapping index 0 (undecided) to `nothing`. +- `seeds_to_integers` looks up each seed value with `indexin`, so every seed + must be complete and every value must be found. +- The result is converted back with `p[c]`, which is where index 0 turns into + the `BoundsError` users see when a constraint defeats the engine. + +The engines share one data structure, `MatrixCoverage` (`coverage_matrix.jl`): +a matrix whose columns are the interactions still to cover, with 0 for +"don't care", and a `remain` counter that partitions covered from uncovered +columns. IPOG (`parameter_order.jl`) extends a design one parameter at a time; +GND (`greedy_tuples.jl`) builds one case at a time from `M` random candidates; +excursions and full factorial enumerate and filter. Mixed strength runs IPOG +once per requested subset, feeding earlier results in as seeds +(`ipog_multi_way`). Constraints are handled by asking the user's predicate at +every choice point whether a partially built case is still allowed, and by +deleting directly forbidden columns from the coverage matrix +(`remove_combinations!`). + +### 2.3 What the documentation tells users + +The docs are honest about the constraint problem +(`docs/src/man/guide.md`, "Exclude forbidden combinations"): the predicate must +tolerate `nothing`, and "the generator *may fail to find a solution* ... you +will see the code try to access a vector at location 0." The JuliaCon paper's +"Statement of Need" and "Comparison of Approaches" sections contain the best +"when to use" writing in the project, but none of it reaches the README or the +docs landing page, which open with the algorithm names. + +## 3. Where it hurts (measured) + +| Probe | Today | +|---|---| +| `all_pairs([1,2,3], ["low","mid","high"], [1.0,3.7,4.9], [:greedy,:relax,:optim])` | returns `Vector{Vector{Any}}`, 10 cases | +| Same, `disallow = (a, b, c) -> a > 2 && b > 2.0` | `MethodError: no method matching isless(::Float64, ::Nothing)` | +| Count calls to a 4-parameter `disallow` during `all_pairs` | 74 calls, 64 of them with at least one `nothing` argument | +| Two constraints whose *combination* makes one pair impossible (Sec. 3.1), IPOG | `BoundsError: attempt to access 2-element Vector{Symbol} at index [0]` | +| Same, GND | never returns (killed after 90 s) | +| Same, `pairs_excursion` | returns 3 cases; the value `:lu` never appears and nothing says so | +| `seeds = [[1, "mid", nothing]]` (partial seed) | `MethodError: Cannot convert an object of type Nothing to Int64` | +| `seeds = [[1, "medium", 3.7]]` (typo in a seed value) | the same `convert` error, no mention of which value or parameter | +| `wayness = Dict(3 => [[3:6], [25:30]])`, copied from `guide.md` | `MethodError: Cannot convert Int64 to UnitRange{Int64}` | +| `wayness = Dict(3 => [3:6, 25:30])` | also a `MethodError`; only `[collect(3:6), collect(25:30)]` works | +| Pass `w = Dict(3 => [[1,3,4]])` as `wayness` | `w` is mutated: afterwards `Dict(2 => [[1,2,3,4]], 3 => [[1,3,4]])` | +| A parameter with one value | `DomainError: Each argument should be a list of parameter values` | +| `n_way = 3` with two parameters | `BoundsError ... at index [1:3]` | + +### 3.1 The constraint problem, precisely + +The running example for this document is a solver configuration: + +```julia +mode = [:fast, :exact] +solver = [:none, :lu, :qr] # :lu and :qr only make sense for :exact +tol = [1e-3, 1e-6] # :exact refuses the loose tolerance +``` + +The two rules are "fast implies no solver" and "exact implies tight tolerance". +Neither rule mentions `(solver, tol)`, but together they make the pair +`(solver = :lu, tol = 1e-3)` impossible: the only mode that allows `:lu` is +`:exact`, and `:exact` forbids `1e-3`. Today's engines remove the *directly* +forbidden pairs from the goal, then try to cover `(:lu, 1e-3)` anyway. IPOG +leaves a 0 it cannot fill and crashes on the way out; GND loops forever looking +for a candidate that scores. Five distinct usability defects sit on top of that +correctness defect: + +1. The predicate takes positional arguments. With 12 parameters, the user + writes `(_, _, m, _, _, s, _...) -> ...` and gets it wrong silently. +2. The predicate receives `nothing` for undecided parameters, so every + comparison other than `==` must be guarded. The docs explain this; users + still hit it first. +3. The engine can only ask "is this partial case still allowed?", which is + a weaker question than "can this partial case still be completed?", and the + gap is exactly the implicit-infeasibility case above. +4. The polarity is inverted from how people state rules ("exact *requires* + tight tolerance"), so users write the negation by hand. +5. Nothing reports what was excluded. In the excursion engine, coverage is + silently lost. + +The literature name for the fix is *forbidden tuples* (Yu, Lei, Kacker, Kuhn, +"Constraint handling in combinatorial test generation using forbidden tuples", +ICSTW 2015; it is how NIST's ACTS handles constraints). The interface below is +designed so that the library owns the constraint as data, which is what makes +the fix possible. + +## 4. Design principles + +- **One concept per level.** A one-line call stays one line. Naming + parameters is one added concept and unlocks everything else. Constraints, + seeds, strength, and inspection each add exactly one more. +- **The library owns constraints as data, not as opaque callables.** It can + then evaluate them safely, detect infeasibility, print them in error messages, + serialize them, and hand them to a solver later. +- **User code is never called with partial information.** No `nothing`, no + `missing`, no index 0. +- **Nothing is dropped silently.** Infeasible interactions, rejected seeds, + and unused values are reported. +- **The result is a first-class Julia value.** Typed, iterable, a table, and + self-describing at the REPL. +- **Guarantees are stated in the docstrings**: every returned case satisfies + every constraint; every feasible t-way interaction is covered; seeds appear + first, in order; IPOG is deterministic. + +## 5. The proposed interface, by level + +### Level 0: one line, positional (unchanged shape, better result) + +```julia +using UnitTestDesign + +cases = all_pairs([1, 2, 3], ["low", "mid", "high"], [1.0, 3.7, 4.9]) +for (n, level, tol) in cases + @test integrate(n, level, tol) ≈ reference(n, level, tol) +end +``` + +`cases` is a `TestCases{Tuple{Int64, String, Float64}}`, an `AbstractVector`. +`case[1]` still works, so existing loops keep working, but the element type is +concrete and `f(case...)` is type-stable. Single-valued parameters are allowed +(a fixed argument is just an argument with one representative). The same +positional form exists for `all_values`, `all_triples`, `covering`, +`full_factorial`, and `excursions`. + +### Level 1: named parameters + +Naming is the one concept that unlocks constraints, partial seeds, subset +strength, tables, and readable test names. Parameters are `name => values` +pairs, in the order the function under test takes them: + +```julia +cases = all_pairs(:n => [1, 2, 3], :level => ["low", "mid", "high"], :tol => [1.0, 3.7, 4.9]) + +@testset "integrate $(case)" for case in cases + (; n, level, tol) = case # property destructuring + @test integrate(n, level, tol) ≈ reference(; case...) # or keyword splat +end +``` + +`cases` is now a `TestCases{@NamedTuple{n::Int64, level::String, tol::Float64}}`. +A `NamedTuple` literal `(n = [1, 2, 3], level = [...])` is accepted as a single +positional argument as well, since it is the natural Julia spelling. + +When the same parameters feed several designs, or need to be inspected, they +live in a `ParameterSpace`: + +```julia +space = ParameterSpace( + :mode => [:fast, :exact], + :solver => [:none, :lu, :qr], + :tol => [1e-3, 1e-6]; + constraints = [ + @forbid(mode == :fast && solver != :none), + @require(mode == :fast || tol < 1e-4), + ], +) +quick = all_pairs(space) # CI +nightly = all_triples(space) # scheduled +``` + +Every design function accepts either a `ParameterSpace` or the pairs directly, +with `constraints =` available in both places. + +### Level 2: constraints + +Three spellings, from most to least common: + +```julia +# 1. A literal forbidden combination. No logic, no macro. +forbid(mode = :fast, solver = :lu) + +# 2. An expression over parameter names. Bare identifiers are parameter names; +# `$x` pulls a variable in from the surrounding scope. +@forbid mode == :fast && solver != :none +@require mode == :fast || tol < $threshold +@forbid solver in (:lu, :qr) && !isfinite(tol) + +# 3. A function, when the logic is too big for an expression. Names are explicit. +forbid(:mode, :solver) do mode, solver + mode == :fast && solver != :none +end +``` + +`@require e` is exactly `@forbid !(e)`; having both verbs keeps the polarity +visible in each rule. All three construct the same `Constraint` value, which +remembers its scope (the parameter names it mentions), its predicate, and its +source text for printing. + +**Semantics.** For each constraint, the library evaluates the predicate once +for every combination of values of the parameters in its scope (the product of +those domains, which is 6 evaluations for `(mode, solver)` above) and records +the combinations that are forbidden. From then on: + +- A complete case is valid iff none of its projections onto a scope is a + forbidden combination. +- A partial case is *rejected* if it already contains a forbidden combination + and *dead* if no valid completion exists; deadness is decided by a + backtracking completion search over the remaining parameters, not by asking + the user. +- The coverage goal is every strength-t interaction that occurs in at least + one valid complete case. Directly forbidden interactions are excluded + trivially; implicitly infeasible ones are found by the completion search and + listed in the result's report. + +The user's predicate is never called during search and never with anything but +values drawn from the declared domains, so `tol < 1e-4` needs no guard. The +cost is the product of the scope's domains; a constraint mentioning ten +four-valued parameters is about a million evaluations, so the library warns +above a threshold and suggests splitting the rule. Real constraints mention +two to four parameters. + +**Errors at construction, not deep in the engine:** + +``` +ArgumentError: @forbid(mode == :fast && solvr != :none) mentions `solvr`, +which is not a parameter. Parameters are (:mode, :solver, :tol). +If `solvr` is a variable, write `$solvr`. +``` + +**Dependent parameters** are the most common real constraint ("`solver` only +applies when `mode == :exact`"). The recipe is a sentinel value plus one +`@require`, which the design handles without special support: + +```julia +:solver => [:none, :lu, :qr] +@require mode == :exact || solver == :none +``` + +**Reporting.** Running the running example gives (prototype output, Appendix B): + +``` +5 test cases · pairwise · 3 parameters · full factorial would be 12 (5 valid) + mode solver tol + :fast :none 0.001 + :exact :none 1.0e-6 + :exact :lu 1.0e-6 + :exact :qr 1.0e-6 + :fast :none 1.0e-6 +Constraints forbid 3 pairs directly and make 2 more unreachable: (solver = :lu, tol = 0.001), (solver = :qr, tol = 0.001). +``` + +The last line is the one that today is a `BoundsError`. It is also useful in +its own right: a user who did not intend `(solver = :lu, tol = 1e-3)` to be +untestable finds out here. + +### Level 3: seeds, strength, engines, and excursions + +**Seeds** are cases that must appear. They are named tuples, may be partial, +and appear first in the output in the order given: + +```julia +all_pairs(space; seeds = [ + (mode = :exact, solver = :lu, tol = 1e-6), # the production configuration + (mode = :fast,), # partial: the engine fills the rest +]) +``` + +A seed that violates a constraint or names a value outside its domain is an +error that names the seed, the parameter, and the offending value. An existing +`TestCases` value is accepted as `seeds`, which is how a suite is extended +(Sec. 6.3). + +**Strength** replaces `n_way`, and `stronger` replaces `wayness`: + +```julia +covering(space; strength = 2) # == all_pairs(space) +covering(space; strength = 2, stronger = [(:mode, :solver, :tol) => 3]) +``` + +Subsets are named, not indexed, so the guide example +`Dict(3 => [[3:6], [25:30]])` becomes `stronger = [(:c, :d, :e, :f) => 3, +(:y, :z) => 3]` and cannot be mistyped into a `convert` error. `strength` +greater than the number of parameters is an `ArgumentError`; `strength` equal +to it is a full factorial. `all_values`, `all_pairs`, and `all_triples` remain +as the names people search for. `all_tuples` remains as a deprecated alias of +`covering`. + +**Engines** keep their literature names but hide their knobs behind readable +ones: + +```julia +covering(space; engine = IPOG()) # default, deterministic +covering(space; engine = GND(rng = Xoshiro(1), candidates = 50)) # `candidates` was `M` +``` + +**Excursions** get an explicit base case. Today the base is silently "the first +value of every parameter", which is the single most surprising fact about +`values_excursion`: + +```julia +excursions(space; from = (mode = :exact, solver = :lu, tol = 1e-6), distance = 1) +excursions(space; distance = 2) # base defaults to first values, and the docstring says so +``` + +`distance` is how many parameters may differ from the base at once, so +`values_excursion`, `pairs_excursion`, and `triples_excursion` become aliases +for distances 1, 2, and 3 and stay exported. Because excursions are filtered +by the same constraint machinery, the report says which values never appear. + +**Full factorial** takes the same space and constraints and reports the count +before and after constraints. + +### Level 4: inspection + +```julia +design_sizes(:n => [1, 2, 3], :level => ["low", "mid", "high"], :tol => [1.0, 3.7, 4.9], :kind => [:greedy, :relax, :optim]) +# strategy cases (measured with today's engines) +# full_factorial 81 (with constraints: total and valid counts) +# all_values 3 +# all_pairs 10 +# all_triples 31 +# excursions(1) 9 +# excursions(2) 33 + +coverage(cases, space; strength = 2) # (covered = 11, feasible = 11, missing = []) +missing_interactions(handwritten, space) # [(solver = :qr, tol = 1.0e-6), ...] +report(cases) # the summary the REPL prints, plus the unreachable list +``` + +`design_sizes` is the "should I bother?" function; it answers the user's +question before they commit. `missing_interactions` is the on-ramp for a +project with an existing suite: convert the existing calls to named tuples, +ask what pairs they miss, and pass them as `seeds` to get the minimal +extension. + +### Reference: proposed signatures + +```julia +ParameterSpace(pairs::Pair{Symbol}...; constraints = Constraint[]) +ParameterSpace(nt::NamedTuple; constraints = Constraint[]) + +covering(space; strength = 2, stronger = Pair[], constraints = [], seeds = [], engine = IPOG()) +all_values(space; kw...) # strength = 1 +all_pairs(space; kw...) # strength = 2 +all_triples(space; kw...) # strength = 3 +excursions(space; from = nothing, distance = 1, constraints = [], seeds = []) +full_factorial(space; constraints = []) +# every function above also accepts `pairs::Pair{Symbol}...` or `values::AbstractVector...` in place of `space` + +forbid(; name = value, ...) -> Constraint +forbid(names::Symbol...) do values... end -> Constraint +require(names::Symbol...) do values... end -> Constraint +@forbid expr; @require expr -> Constraint + +design_sizes(space; strengths = 1:3, distances = 1:2) +coverage(cases, space; strength = 2) +missing_interactions(cases, space; strength = 2) +report(cases) + +struct TestCases{T} <: AbstractVector{T} # T is a Tuple or NamedTuple type +IPOG(); GND(; rng = Random.default_rng(), candidates = 50) +``` + +Removed from the public surface: `generate_tuples`, `Excursion`, `Counter`, +`n_way`, `wayness`, `disallow` (kept for one minor version as a deprecated +keyword that wraps the function form with `nothing`-safe semantics, see Sec. 9). + +## 6. The result type + +### 6.1 `TestCases` + +```julia +struct TestCases{T} <: AbstractVector{T} + cases::Vector{T} + space::ParameterSpace + strategy::Symbol # :covering, :excursion, :full_factorial + strength::Int # or distance, for excursions + seeded::Int # how many leading cases came from seeds + infeasible::Vector{NamedTuple} # interactions dropped from the goal +end +``` + +It is a vector, so everything that works on a vector works on it. Because it +is an `AbstractVector` of `NamedTuple`s it is already a Tables.jl row table, +so `DataFrame(cases)` and `CSV.write("cases.csv", cases)` work with no +dependency added here. `show` prints the one-line summary and an aligned +table, truncated like a `DataFrame`. The metadata is what `report`, +`coverage`, and the deprecation shims read. + +### 6.2 REPL output as documentation + +``` +julia> all_pairs(:n => [1, 2, 3], :level => ["low", "mid", "high"], :tol => [1.0, 3.7, 4.9], :kind => [:greedy, :relax, :optim]) +10 test cases · pairwise · 4 parameters · full factorial would be 81 + n level tol kind + 1 "low" 1.0 :greedy + 2 "mid" 3.7 :relax + ⋮ +``` + +"10 of 81" is the whole pitch of the package, and it should be the first thing +a user sees. + +### 6.3 Extending an existing suite + +```julia +existing = [(mode = :fast, solver = :none, tol = 1e-3), (mode = :exact, solver = :lu, tol = 1e-6)] +missing_interactions(existing, space) # what the hand-written tests miss +all_pairs(space; seeds = existing) # the smallest set that keeps them and fixes the gaps +``` + +## 7. Positioning: telling people when to use it + +### 7.1 The decision in one table + +This belongs at the top of the README and the docs landing page, before any +function name. + +| Your situation | Reach for | +|---|---| +| You can list a few representative values per argument, the function has several arguments, and you suspect bugs live in *combinations* of arguments (an `if` on one option inside a branch on another). | A covering design: `all_pairs`, `all_triples`. | +| Same, and each run is cheap and the product is small. | `full_factorial` or `Iterators.product`. | +| You have one known-good configuration and want to know which single (or paired) changes break it. | `excursions`. | +| You can *generate* values but not enumerate representatives (strings, trees, arbitrary floats), runs are cheap, and you want failing inputs shrunk. | Property-based testing (Supposition.jl, PropCheck.jl). | +| You already have hand-written cases and want to know what interactions they miss. | `missing_interactions`, then `seeds`. | +| Both apply. | Use covering designs to pick *which equivalence class* each argument draws from, and a generator to draw within it. | + +### 7.2 The one-sentence explanation + +Anchor to something users already know: `Iterators.product(a, b, c)` is every +combination; `all_pairs(a, b, c)` is the smallest subset of it in which every +pair of values still occurs together. Random sampling gets there too, but +slowly and without a guarantee (measured, uniform random cases, 200 trials): + +| parameters × values | designed pairwise cases | random cases needed to cover every pair, median [10%, 90%] | pairs covered by that many random cases | +|---|---|---|---| +| 5 × 4 | 16 | 84 [65, 116] | 64% | +| 10 × 4 | 28 | 106 [87, 131] | 84% | +| 10 × 8 | 113 | 530 [461, 633] | 83% | +| 40 × 4 | 45 | 151 [133, 183] | 95% | +| 40 × 8 | 166 | 709 [632, 845] | 93% | + +The honest framing: random testing spends four to six times as many runs to +reach the same pairwise guarantee, and property-based testing spends them to +buy something different (generation and shrinking). Designed cases win when +runs are expensive or the argument space is large, and lose when the user +cannot name representatives. The paper already says this; the README should. + +### 7.3 Docs restructure + +1. **Home**: the table above, the one-sentence explanation, the 10-of-81 + example with its REPL output, and the PBT bridge in three lines. +2. **Tutorial** (replaces Guide): levels 0 through 4 in order, each a runnable + `@example` block, ending with "extend an existing suite". +3. **Constraints**: the three spellings, the semantics paragraph, the + dependent-parameter recipe, and what the report means. +4. **Choosing strength and strategy**: the `design_sizes` table for a real + example, when to raise strength, when excursions beat coverings. +5. **Engines and algorithms**: the existing IPOG and greedy pages, with the + forbidden-tuple and completion-search additions. +6. **Reference**. + +The paper's "How to Use" section (equivalence classes, oracles, invariants) +belongs in the tutorial, not only in the PDF. + +## 8. What the interface assumes of the engines + +The interface is only honest if the engines deliver these. Each is stated as a +contract, with what it costs to meet it. The prototype in Appendix B meets the +first three on top of the *current* IPOG code, which is the evidence that the +list is realistic rather than aspirational. + +- **E1. Constraints arrive as forbidden-combination tables, never as + callables.** A preprocessing step tabulates each `Constraint` over its scope. + The engines' existing integer-vector predicate is kept internally but is now + generated by the library and is partial-safe by construction (a table only + fires once its whole scope is assigned). No engine change. + +- **E2. Implicitly infeasible interactions are detected and dropped from the + goal.** Before generation, each target interaction is checked for a valid + completion by backtracking with forward checking over the remaining + parameters. Cost is negligible with few constraints: the 12-parameter, + four-constraint example in Appendix B takes 0.7 s end to end, most of it in + IPOG. Later optimization: derive minimal forbidden tuples once (Yu et al. + 2015) so per-interaction search is rarely needed. Prototype: done on top of + `ipog_multi` by adding the dead interactions to the predicate. + +- **E3. Engines never leave an unfillable slot.** Replace the greedy + `fill_remaining_missing_values_filter!` and `choose_last_parameter_filter!` + fallbacks with backtracking completion. With E2 in place, this only triggers + for higher-order implicit infeasibility (a partial case whose every pair is + feasible but whose triple is not), which is rare and cheap. This is what + makes "index 0" impossible by construction rather than by luck. + +- **E4. GND cannot spin.** Its candidate loop (`while trial_cnt < M`) must + cap attempts and fall back to backtracking completion; with E2 removing + unreachable goals it also terminates by the same argument as today. + +- **E5. Partial seeds and single-valued parameters.** Measured: `ipog`, + `ipog_multi`, and the greedy generator all accept an arity of 1 and reach + full coverage, and `ipog_multi` given the partial seed `[1, 0, 2]` fills the + zero and keeps the seed first. Only the translation layer + (`seeds_to_integers`, the `< 2` checks in `all_tuples`) forbids them. Small + change. + +- **E6. Coverage is reportable.** `test_coverage` and `tuples_in_trials` + already exist in `coverage_set.jl`; they need to be exposed through + `coverage` and `missing_interactions` in value space rather than index space. + +- **E7. Determinism and reproducibility as stated.** IPOG is deterministic + today. Constraint tabulation and dead-interaction detection must iterate in + a fixed order so the report is stable across runs. + +- **E8. Room to grow, without changing the interface.** Because a + `Constraint` keeps its expression, a package extension can translate the + comparison-and-boolean subset to Z3 (the sketch in `z3_example.jl`) or a SAT + solver for large scopes, and the same strings can be read from TOML for the + command-line branch (`feature/command-line`), which today has no way to + express `disallow` at all. Neither is needed for the first release. + +Suggested order: E1, E2, E5, E6 (the translation layer, no engine surgery), +then E3 and E4 (engine surgery), then E7 checks, then E8 as separate projects. + +## 9. Migration + +| Today | Proposed | Notes | +|---|---|---| +| `all_pairs(v1, v2, v3)` returning `Vector{Vector{Any}}` | same call, returns `TestCases{Tuple{...}}` | breaking only for code that mutates result rows or depends on `Vector{Vector}` | +| `all_tuples(...; n_way = k)` | `covering(...; strength = k)` | `all_tuples` and `n_way` deprecated one minor version | +| `disallow = (a, b, c) -> ...` | `constraints = [@forbid ...]` | `disallow` kept one version: wrapped as a full-scope function-form constraint, so it is evaluated only on complete cases and the `nothing` contract disappears | +| `wayness = Dict(3 => [[3, 4, 5, 6]])` | `stronger = [(:c, :d, :e, :f) => 3]` | positional call: `stronger = [(3, 4, 5, 6) => 3]` accepted | +| `seeds = [[1, "mid", 3.7, :relax]]` | `seeds = [(n = 1, level = "mid", tol = 3.7, kind = :relax)]` | positional call still accepts vectors/tuples; partial seeds new | +| `values_excursion(...)` | `excursions(...; distance = 1)` | old names stay as aliases; `from =` new | +| `engine = GND(M = 50)` | `engine = GND(candidates = 50)` | `M` deprecated | +| `Counter = Int8` | removed | choose internally from arity | +| `generate_tuples`, `Excursion` | unexported | | + +Version: this is a 1.0 candidate. The return-type change is the only break +that cannot be shimmed, and it is the change most worth making. + +## 10. Open questions for the author + +1. **Names.** `ParameterSpace`, `covering`, `stronger`, `forbid`/`require`, + `excursions(; from, distance)`. Alternatives considered: `Factors`/`levels` + (DoE vocabulary, foreign to unit testers), `design` (collides with user + variables), `constraints = [...]` versus `where = ...` (kept `constraints`, + the term every other CIT tool uses). +2. **Should the positional form support constraints at all?** This proposal + says no: constraints need names. A positional user adds names when they + add a constraint, which is one honest step up. +3. **Labels for values.** `:x => ["tiny" => 1e-9, "huge" => 1e9]` would make + test names and reports readable for float and struct values. It adds a + concept; deferred. +4. **Keep GND?** It gives the same guarantees, is slower, and is less + predictable. It sometimes finds smaller designs, which matters when each run + costs hours. Keep, but do not mention it before the "Engines" page. +5. **Generators as values.** The PBT bridge could be a helper that draws from + generator values at iteration time. Values that are functions are also + legitimate arguments to test (`[sin, cos]`), so nothing automatic; document + the pattern instead. +6. **Command line.** With constraints expressible as strings of the safe + subset, the TOML format from `feature/command-line` becomes complete. + Worth reviving after 1.0. + +## Appendix A. Probe results + +Script: `probe.jl` (interface probes), `probe_gnd.jl` (GND hang under a 90 s +alarm), `random_vs_design.jl` (Sec. 7.2 table), run with +`julia --project=.