Repository navigation
Phase 3: engines that keep the promise - #54
Merged
Merged
Conversation
request.jl is the second index-space boundary: a Request validates strength, stronger groups, and must-include rows against the space, exposes dead() and witness() to the engines, classifies every target with attribution, and validate_design certifies a result. GND now takes a fixed default seed and candidates, with M deprecated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…llow generate(::IPOG, request) classifies every target, covers exactly the required ones, and certifies the result. The three IPOG sites use dead() instead of disallow, and ipog_multi_way keeps rows full width so a value set for one stronger group is visible to the next, which fixes invalid rows with overlapping groups. The positional interface routes through Request; disallow, Counter, generate_tuples, and the nothing sentinel are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…layer GND builds each case from candidates checked with dead(), and when no candidate covers anything it builds the next case from the first uncovered target, so it always terminates. Excursions take an explicit base, drop and report dead rows, and reject a forbidden base by rule. Full factorial refuses a product above the limit before enumerating and retains only accepted rows through an in-place odometer. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both engines pass 500 pairwise and 500 three-way random problems, each design certified by the independent checker, and every Phase 3 fixture test is live. A limited search raises ResourceLimitError during classification or placement and never returns an incomplete design. benchmark/run.jl records the baseline against 0.4 in design/benchmark_results_phase3.md. Docs examples that passed disallow are plain code blocks until Phase 7. Closes #51. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
per-request lazy memo, index-space validation, CI benchmark job Excursions take one distance (zero allowed) and reject stronger groups. Excluded records which limit stopped an explanation and an infeasible must-include row's error carries its rules. A lazy rule's memo now lives in the operation's Feasibility, so a TestSpace retains nothing from generation. validate_design checks index rows directly, making full factorial on bench12 about 10x faster. CI gains a benchmark job. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Owner
Author
|
Review round 1 applied (see "Review round 1" and "CI-runner baseline" in
Full suite: 270,373 pass, 13 pending, 0 failures. Matrix green on 1.10 / 1.11 / 1 across ubuntu, macOS, Windows. 🤖 Generated with Claude Code |
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
covering, all_values, all_pairs, all_triples, excursions, and
full_factorial accept a TestSpace, a NamedTuple, pairs, or positional
vectors and return TestCases{T}, a read-only vector of typed rows that
records the space, strategy, engine and seed, must-include count, covered
and excluded targets with attribution. show prints only recorded
bookkeeping. all_tuples, n_way, seeds, wayness, the *_excursion names,
and GND(M=) warn through depwarn; Excursion and Counter are gone.
Excursion notes are values, not positions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
exclusion wording, budget docs, strength-0 semantics for Phase 5 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
field types from values, min(2, n) report default, duplicate wording Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
coverage measures any supplied rows against a space, independent of generator bookkeeping: covered targets need no search, missing ones are classified with attribution, unknown ones are listed without a percentage, ordinary and negative parts are kept separate, and rejected rows contribute nothing. Unconstrained requests no longer materialize their target list for the certification recount. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
report verifies a result independently of its bookkeeping: a guarantee line, the excluded list with attribution, bonus coverage at strength + 1, a prefix curve computed in one pass, and the seed, never printing a percentage while a target is unresolved. design_sizes runs each strategy and tabulates cases, share of the valid product, and pairs and triples covered, recording a resource-limit status instead of a count. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The excursion guarantee exempts must-include rows; report shows its own freshly verified exclusions and labels recorded proofs as a fallback; coverage(cases; strength) keeps stored stronger groups unless replaced; plain() gives a Report, Coverage, or DesignSizes representation with no executable state, tested by round trip; measurement docstring examples are doctests with real output. The release is 0.5.0: the maintainer wants a 0.x version to use before calling anything 1.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…matrix Negative targets are classified through the negative-row search and covered by a sub-request over the other parameters with the same engine, with separate bookkeeping and a ! marker on negative rows. Must-include, full factorial, and excursions accept wrappers under their row policies. realize draws each Partition once under the caller's rng. diagnose ranks combinations present only in failing cases and groups indistinguishable ones; followups finds isolating cases or proves there are none. github_matrix writes a validated GitHub Actions include matrix with JSON.jl. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
figures, JSON field-name validation Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The docs environment pinned Documenter 0.25, which needs JSON 0.21, while the package now requires JSON 1, so CI could not resolve it. Documenter 1 makes missing docstrings and unresolved references errors by default; they stay warnings until Phase 7 completes the manual. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
README and home page lead with the promise, the decision table, and one example with its printed output. The manual is a tutorial in five levels, nine how-to guides by job, explanation pages on values and oracles, interaction coverage and its evidence, constraints, and the engines, a reference organized by purpose with every exported docstring opening on the situation it serves, a 0.4 to 0.5 migration table, developer pages, and a one-pager for AI agents. Doctests run in the docs build and in the test suite; the build fails on any warning. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
wording, two-check lockfile pattern, badge and contract wording Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Phase 7: documentation and first contact
Phase 6: Invalid, Partition, diagnose, github_matrix
Phase 5: measure and explain
Phase 4: the public interface and the result type
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 3 of
design/20260926_implementation_plan.md, stacked on #53 (retarget after it merges). Review note with acceptance results, benchmark tables, and open decisions:design/phase3_review.md.Request/Designlayer: every target classified with attribution,dead()as the only feasibility question an engine asks, final validation certifies every result (src/request.jl).ipog_multi_wayrewritten full-width, fixing a 0.4 bug where overlapping stronger groups produced invalid rows.candidatesin place ofM.disallow,Counter,generate_tuples, and thenothingsentinel are gone; the positional API routes throughRequest.bench12now runs where 0.4 crashed or hung.Full suite: 271,840 pass, 13 broken/skipped (Phases 4–6), 0 failures. Aqua and docs build green.
🤖 Generated with Claude Code