Repository navigation
Phase 1: specification, checker, and scaffolding - #52
Merged
Merged
Conversation
Bump to 1.0.0-DEV with julia = "1.10" and a CI matrix of 1.10, 1.11, and the latest release. Ignore .DS_Store, Manifest.toml, and .vscode/. Move the interface proposals, synthesis, answers, and implementation plan into design/, and add the Phase 3 benchmark procedure, the dependents check with a changelog skeleton, and the Phase 1 review note. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs/src/dev/contract.md states the six-point contract, value identity, resource limits, partitions, negative rows and targets, strategy boundaries, size claims, determinism, must-include, stronger groups, constraint forms, the vocabulary, and the not-now list, in numbered clauses that later phases and tests cite. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
test/checker.jl is a brute-force, value-space oracle that shares no code with src/. test/random_problems.jl draws constrained problems and runs both engines through the 0.4 API; IPOG's BoundsError on infeasible targets is the one expected failure (@test_broken) and GND is skipped on problems with implied targets (@test_skip) until Phase 3 closes #51. test/fixtures.jl holds named fixtures with pending engine tests marked by their activation phase. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eze fixtures GND attempt-cap errors are ordinary failures and IPOG crashes get a ceiling, so only the known BoundsError stays @test_broken. Contract 1.24 now speaks of ordinary rows only, since negative rows can be valid in a space with no valid ordinary row; 2.3 keeps identity for stored choices and returned cases while naming the permitted output projections. The checker describes "implied" as infeasible with no direct rule match. The greedy dead ends from the random run, the empty-ordinary negative-seed case, and a bench12 replacement for Fable's 12-parameter example are frozen as fixtures with their legacy baseline recorded, and UNITTESTDESIGN_TEST_LONGER lets a CI job run the full random gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Owner
Author
|
Review round 1 applied (see the "Review round 1" section of Full suite: 230,617 pass, 43 broken/skipped (all deliberate), 0 failures. 🤖 Generated with Claude Code |
…lity Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…abulation The model layer of the 1.0 contract: a TestSpace validates names and domains by value identity, holds Constraints built from patterns, named predicates, whole-case predicates, or @forbid/@require, and tabulates each rule into a RuleTable over ordinary values, lazily above tabulation_limit. Negative-row rule selection is active_tables. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tial Backtracking with forward checking over connected components, per-component witness caches, a node budget that returns unknown and never memoizes an exhausted search, deletion search for implied exclusions with a separate explanation budget, and ResourceLimitError. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
src/explain.jl builds the ordinary or negative-row search for an assignment, maps rule numbers back to constraint positions, and prints one sentence per outcome. The checker adapter test_space moves into the shared Checker test module, the 12 Phase 2 pending lines become real tests, and test_explain.jl cross-checks classify against the checker on every fixture and 100 random problems. to_indices is renamed case_indices to avoid shadowing Base. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…emantics Legacy generation rejects Invalid and Partition values until Phase 6, and two-domain full_factorial no longer hits a MethodError. The macro walker scopes ->, let, generators, comprehensions, and do blocks with Julia's rules and rejects other binding forms with a pointer to the function form. The contract states that the node budget does not count rule checks, results report nodes and checks, and a lazy rule's memo is documented as space-owned tabulation with memo_size to measure it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
request.jl is the second index-space boundary: a Request validates strength, stronger groups, and must-include rows against the space, exposes dead() and witness() to the engines, classifies every target with attribution, and validate_design certifies a result. GND now takes a fixed default seed and candidates, with M deprecated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…llow generate(::IPOG, request) classifies every target, covers exactly the required ones, and certifies the result. The three IPOG sites use dead() instead of disallow, and ipog_multi_way keeps rows full width so a value set for one stronger group is visible to the next, which fixes invalid rows with overlapping groups. The positional interface routes through Request; disallow, Counter, generate_tuples, and the nothing sentinel are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…layer GND builds each case from candidates checked with dead(), and when no candidate covers anything it builds the next case from the first uncovered target, so it always terminates. Excursions take an explicit base, drop and report dead rows, and reject a forbidden base by rule. Full factorial refuses a product above the limit before enumerating and retains only accepted rows through an in-place odometer. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both engines pass 500 pairwise and 500 three-way random problems, each design certified by the independent checker, and every Phase 3 fixture test is live. A limited search raises ResourceLimitError during classification or placement and never returns an incomplete design. benchmark/run.jl records the baseline against 0.4 in design/benchmark_results_phase3.md. Docs examples that passed disallow are plain code blocks until Phase 7. Closes #51. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
per-request lazy memo, index-space validation, CI benchmark job Excursions take one distance (zero allowed) and reject stronger groups. Excluded records which limit stopped an explanation and an infeasible must-include row's error carries its rules. A lazy rule's memo now lives in the operation's Feasibility, so a TestSpace retains nothing from generation. validate_design checks index rows directly, making full factorial on bench12 about 10x faster. CI gains a benchmark job. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
covering, all_values, all_pairs, all_triples, excursions, and
full_factorial accept a TestSpace, a NamedTuple, pairs, or positional
vectors and return TestCases{T}, a read-only vector of typed rows that
records the space, strategy, engine and seed, must-include count, covered
and excluded targets with attribution. show prints only recorded
bookkeeping. all_tuples, n_way, seeds, wayness, the *_excursion names,
and GND(M=) warn through depwarn; Excursion and Counter are gone.
Excursion notes are values, not positions.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
exclusion wording, budget docs, strength-0 semantics for Phase 5 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
field types from values, min(2, n) report default, duplicate wording Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
coverage measures any supplied rows against a space, independent of generator bookkeeping: covered targets need no search, missing ones are classified with attribution, unknown ones are listed without a percentage, ordinary and negative parts are kept separate, and rejected rows contribute nothing. Unconstrained requests no longer materialize their target list for the certification recount. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
report verifies a result independently of its bookkeeping: a guarantee line, the excluded list with attribution, bonus coverage at strength + 1, a prefix curve computed in one pass, and the seed, never printing a percentage while a target is unresolved. design_sizes runs each strategy and tabulates cases, share of the valid product, and pairs and triples covered, recording a resource-limit status instead of a count. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The excursion guarantee exempts must-include rows; report shows its own freshly verified exclusions and labels recorded proofs as a fallback; coverage(cases; strength) keeps stored stronger groups unless replaced; plain() gives a Report, Coverage, or DesignSizes representation with no executable state, tested by round trip; measurement docstring examples are doctests with real output. The release is 0.5.0: the maintainer wants a 0.x version to use before calling anything 1.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…matrix Negative targets are classified through the negative-row search and covered by a sub-request over the other parameters with the same engine, with separate bookkeeping and a ! marker on negative rows. Must-include, full factorial, and excursions accept wrappers under their row policies. realize draws each Partition once under the caller's rng. diagnose ranks combinations present only in failing cases and groups indistinguishable ones; followups finds isolating cases or proves there are none. github_matrix writes a validated GitHub Actions include matrix with JSON.jl. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
figures, JSON field-name validation Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The docs environment pinned Documenter 0.25, which needs JSON 0.21, while the package now requires JSON 1, so CI could not resolve it. Documenter 1 makes missing docstrings and unresolved references errors by default; they stay warnings until Phase 7 completes the manual. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
README and home page lead with the promise, the decision table, and one example with its printed output. The manual is a tutorial in five levels, nine how-to guides by job, explanation pages on values and oracles, interaction coverage and its evidence, constraints, and the engines, a reference organized by purpose with every exported docstring opening on the situation it serves, a 0.4 to 0.5 migration table, developer pages, and a one-pager for AI agents. Doctests run in the docs build and in the test suite; the build fails on any warning. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
wording, two-check lockfile pattern, badge and contract wording Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Phase 7: documentation and first contact
Phase 6: Invalid, Partition, diagnose, github_matrix
Phase 5: measure and explain
Phase 4: the public interface and the result type
Phase 3: engines that keep the promise
Phase 2: the model: TestSpace, rules, feasibility
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 1 of
design/20260926_implementation_plan.md. Review note with artifacts, numbers, pending tests, and open decisions:design/phase1_review.md.docs/src/dev/contract.md, 168 numbered clauses (docs build clean).test/checker.jl, verified against hand-enumerated fixtures (221 assertions), including the solver example's 3 direct and 2 implied exclusions.BoundsErroron infeasible targets is the one@test_broken; GND is@test_skipped on problems with implied targets (it hangs) until Phase 3 closes Implicit constraints crash IPOG and hang GND #51.1.0.0-DEV,julia = "1.10", CI on 1.10 / 1.11 / 1, design files moved todesign/.Full suite: 229,311 pass, 34 broken/skipped (all deliberate), 0 failures.
🤖 Generated with Claude Code