Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion .claude/agents/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,14 +10,19 @@ Sub-agents the main Claude Code agent can delegate to. Each one runs in its own

The kit's bar for shipping a sub-agent: it has to do something a skill can't. In practice that means either **tool restriction** (the sub-agent literally can't call `Edit` / `Write`) or **context isolation** for work that would otherwise dump thousands of lines into the main agent's window.

A "specialist who writes blocks" or "specialist who writes REST endpoints" doesn't clear that bar — it's a skill in disguise. So the kit ships exactly two sub-agents, both read-only audits.
A "specialist who writes blocks" or "specialist who writes REST endpoints" doesn't clear that bar — it's a skill in disguise. Nor does "a planning agent" or "a coding agent": planning is already a skill, and a coding agent needs the main thread's context and hands back a diff rather than a conclusion. So the kit ships three sub-agents, all read-only, each covering one thing the main agent can't do safely or cheaply in its own window.

## Shipped sub-agents

| Sub-agent | Tools | When to invoke |
|---|---|---|
| [`plan-reviewer`](./plan-reviewer.md) | Read, Grep, Glob, Bash | During Phase 2.5 of `wordpress-feature`, after `scripts/open-plan-pr.sh` opens the plan PR. Audits spec + plan against PLANNING.md and the constitution. |
| [`security-reviewer`](./security-reviewer.md) | Read, Grep, Glob, Bash | Before merging a feature PR, before a release, after any change to input handling. |
| [`playground-verifier`](./playground-verifier.md) | Read, Grep, Glob, Bash, `wp-playground` MCP | Before merging a feature PR, and right after a scaffold. Boots the plugin in a real WordPress instance and verifies it activates, serves its routes, and uninstalls cleanly. |

The three map onto what can go wrong at three different times: the plan is wrong (`plan-reviewer`), the code is unsafe (`security-reviewer`), or the code is fine on paper and broken in practice (`playground-verifier`). The last one is the only gate in the kit that boots WordPress — everything else, `quality.sh` included, is static.

`security-reviewer` and `playground-verifier` read the same diff and don't depend on each other, so dispatch them in the same message and let them run concurrently.

Block and REST work is handled by the `wp-block-development` and `wp-rest-api` skills the kit pulls from [WordPress/agent-skills](https://github.com/WordPress/agent-skills) into `.claude/skills/` — those run on the main agent's context, no handoff needed.

Expand Down
233 changes: 233 additions & 0 deletions .claude/agents/playground-verifier.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .claude/skills/wordpress-feature/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Decide first: **feature** (gets spec + plan, possibly its own PR) or **maintenan
- The user can override either direction.
5. **Implement.** After the plan PR merges (or immediately for "proceed"), mirror open plan steps into the harness task tool (e.g. `TaskCreate` if available). Tick off as you go; append to `progress.md`.
6. **Findings (lazy).** When an MCP lookup (`wp-devdocs`, `wp-blockmarkup`) surfaces a non-obvious fact that shaped the code, append to `findings.md`. Skip if nothing surprised you.
7. **Ship.** Run `./scripts/quality.sh`. Open the feature PR. Dispatch `security-reviewer` on the diff. After merge, archive the feature dir to `.claude/plans/archive/{YYYY-MM-DD}-{slug}/`.
7. **Ship.** Run `./scripts/quality.sh`. Open the feature PR. Dispatch `security-reviewer` and `playground-verifier` on the diff — in one message so they run concurrently; they don't depend on each other. `quality.sh` proves the code is well-formed; `playground-verifier` proves it runs. After merge, archive the feature dir to `.claude/plans/archive/{YYYY-MM-DD}-{slug}/`.

## Maintenance mode

Expand Down
4 changes: 2 additions & 2 deletions .claude/skills/wordpress-scaffold/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ Scaffolds a brand-new WordPress plugin through a short interview. **Greenfield o

5. **Generate.** Standard plugin layout: `{slug}.php` (bootstrap with header, `*_VERSION`/`*_PATH`/`*_URL`/`*_FILE` constants, single `get_instance()` class, top-level activation/deactivation hooks), `includes/class-utils.php`, `includes/class-settings.php` (if selected), `includes/api/class-rest.php` + `class-rest-example.php` (if selected), `src/blocks/test-block/` (if selected), `src/extensions/` (if selected), `languages/`, `package.json` (if blocks/extensions), `composer.json`, `phpcs.xml.dist`, `readme.txt`, `uninstall.php`, `.gitignore`. Every PHP file MUST start with the `ABSPATH` guard.

6. **Deliver.** Show the file tree. Print the three install commands (`composer install`, `npm install` if blocks/extensions, `npm run build` if blocks/extensions). Write `progress.md` with `status: complete, last_completed: scaffold` and immediately move `features/000-initial-scaffold/` to `.claude/plans/archive/{YYYY-MM-DD}-000-initial-scaffold/`. Offer to spin up `wp-playground` for a smoke test. Tell the user: **next features go through `wordpress-feature`.**
6. **Deliver.** Show the file tree. Print the three install commands (`composer install`, `npm install` if blocks/extensions, `npm run build` if blocks/extensions). Write `progress.md` with `status: complete, last_completed: scaffold` and immediately move `features/000-initial-scaffold/` to `.claude/plans/archive/{YYYY-MM-DD}-000-initial-scaffold/`. Then dispatch `playground-verifier` — a scaffold that doesn't activate is worse than no scaffold, and nothing static catches it. Tell the user: **next features go through `wordpress-feature`.**

## Hard rules (this skill)

Expand All @@ -45,7 +45,7 @@ Every PHP file MUST start with `if ( ! defined( 'ABSPATH' ) ) { exit; }`. Every

## After generation

Remind the user to run `./vendor/bin/phpcs`, verify activation → deactivation → uninstall leaves no orphan data, commit `composer.lock` + `package-lock.json`, and run `security-reviewer` before the first PR.
Remind the user to run `./vendor/bin/phpcs`, commit `composer.lock` + `package-lock.json`, and run `security-reviewer` before the first PR. The activation → deactivation → uninstall cycle is `playground-verifier`'s job — dispatch it rather than asking the user to check by hand.

## References

Expand Down
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,28 @@ The kit's evolution, kept for humans. Per-feature progress lives in `.claude/pla

## v1.0.4 — 2026-09-16

### Added

- **`playground-verifier` sub-agent** (`.claude/agents/playground-verifier.md`). The kit's first *runtime* gate. Boots the plugin in an ephemeral WordPress Playground instance — at the WP and PHP versions the plugin's own header declares, not at latest — mounts the working tree, then activates, reads the error log, confirms every `register_rest_route` call actually resolves, checks that mutating routes reject unauthenticated requests, deactivates, uninstalls, and verifies no options are left behind. Read-only plus the `wp-playground` MCP tools: no `Edit`, no `Write`. Always stops the instance, including on early exit.

This closes the hole that v1.0.2 exposed. That release fixed a scaffold which fataled the instant it was activated, and the fix note said it plainly: *"None of the static gates caught it, because nothing in `quality.sh` boots WordPress."* phpcs reads syntax, `security-reviewer` reads patterns, `plan-reviewer` reads markdown. Nothing ran the plugin. Now something does.

The definition was then **corrected against a live run** before shipping. Dispatching it at the kit's own `pl-example` template surfaced 13 defects in its own playbook, two of which would have produced false passes: `get_playground_logs` surfaces the Node process's output and **not** PHP errors (those land in `wp-content/debug.log`, invisible to it — an agent following the original instructions would report "no errors" on a fataling plugin), and the unauthenticated-REST check could not run at all, because Playground's auto-login mu-plugin 302-redirects every cookie-less request until curl dies at exit 47. The playbook now names `debug.log` as the evidence channel, requires a deliberate `trigger_error` probe to prove logging is on before an empty log means anything, carries working recipes for the commands the MCP bridge doesn't implement (`wp plugin uninstall`, `is_plugin_active()`, `option get`'s missing-vs-empty ambiguity), confirms the booted versions from inside the instance rather than trusting the CLI banner (which reported PHP 8.3 / WP latest for an instance serving 8.2.33 / 6.7.7), and filters the route dump to the plugin's namespace instead of flooding context with every core route.

It also gained one deliberate carve-out from "never edit": VFS-only scratch copies at non-mounted paths, so the agent can remove a dependency to prove a fix carries its own weight — which is how it established that the v1.0.2 autoload fix is real and not merely masked by Composer's classmap. Host tree untouched, verified with `git status --porcelain`.

- **Sub-agent contract tests** (`tests/agents/test-agent-contracts.sh`, 43 assertions). The read-only guarantee lives in each agent's `tools:` allowlist, not in its prompt body — a prompt saying "you do not modify code" is advisory, an omitted `Edit` is structural. These tests assert the structural half: no kit sub-agent declares `Edit` / `Write` / `MultiEdit` / `NotebookEdit`, every `name:` matches its filename, every agent has a `description:` (the main agent matches intent against it), `playground-verifier` holds the four `wp-playground` tools it can't work without, and no skill dispatches an agent that isn't on disk. Six further assertions pin the false-pass traps above, so a future edit that drops the `debug.log` warning or the auto-login recipe reintroduces a gate that cannot fail. Suite total: 92 → 135.

### Changed

- **`wordpress-feature` step 7** now dispatches `security-reviewer` and `playground-verifier` together, in one message, so they run concurrently against the same diff. `quality.sh` proves the code is well-formed; `playground-verifier` proves it runs.
- **`wordpress-scaffold` step 6** dispatches `playground-verifier` instead of offering a manual Playground smoke test, and its "After generation" reminder hands the activation → deactivation → uninstall cycle to the agent rather than to the user. A scaffold that doesn't activate is worse than no scaffold.
- **`docs/06-sub-agents.md`** gains a "why not planner / coder / tester" section. One-agent-per-phase is the structure most people reach for first; the doc now runs each candidate through the kit's three criteria (context isolation, tool discipline, parallelism) and shows why only the verifier clears the bar. The rule it lands on: delegate work whose **output is a conclusion**, not work whose output is a diff.

### Fixed

- **`playground-verifier`'s own playbook**, corrected against a second live run: `WP_DEBUG_DISPLAY` is false by default in Playground, which made the "notices in the page body" check structurally incapable of failing; the `wp-admin/includes/` caveat was written as a quirk of `is_plugin_active()` when it applies to `activate_plugin()` and `get_plugins()` equally; there was no recipe for *authenticated* admin loads (the inverse of the anonymous one — omit the suppression cookie and let auto-login work); and the deactivation criterion put transients at High, which flags correct conventional behaviour on nearly every plugin. The fresh-clone reproduction now uses `git archive HEAD` rather than a hand-derived exclusion list that drifts from `.gitignore`.

- **The planning layer was dead on Linux.** `session-start.sh`, `user-prompt-submit.sh` and `stop.sh` each derived a file's mtime with `stat -f %m "$f" || stat -c %Y "$f"`. That is correct on BSD/macOS. On GNU/Linux `-f` means *file system status*, so `%m` is not a format string — it is parsed as a second FILE argument. GNU prints a six-line filesystem dump for `$f`, **then** exits 1 because no file named `%m` exists, so the `||` fallback runs too and the capture holds both. Arithmetic on that blob is a syntax error, the comparison silently evaluates false, and all three hooks concluded there was no active feature.

Net effect: no plan injection on every turn, no session-start banner, no `progress.md` timestamping — the entire Descrição mechanism, silently inert, on every Linux machine. Green on the author's laptop the whole time.
Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Senior WordPress developers don't need to write boilerplate anymore. What we sti
|---|---|---|
| **Delegação** | [`.claude/skills/`](./.claude/skills/) | Interview-driven workflows (`wordpress-scaffold`, `wordpress-feature`) that hand structured work to the agent |
| **Descrição** | [`.claude/plans/`](./.claude/plans/) + `CLAUDE.md` / `AGENTS.md` | Durable description the agent reads first — constitution, spec, plan, progress |
| **Discernimento** | [`.claude/agents/`](./.claude/agents/) | Read-only sub-agents (`plan-reviewer`, `security-reviewer`) — independent judgement before sign-off |
| **Discernimento** | [`.claude/agents/`](./.claude/agents/) | Read-only sub-agents (`plan-reviewer`, `security-reviewer`, `playground-verifier`) — independent judgement before sign-off |
| **Diligência** | [`.claude/hooks/`](./.claude/hooks/) | Deterministic gates: pre-commit quality, post-edit lint, plan injection, progress timestamping |

Plus curated MCP servers and the scaffolding CLI that sets it all up in one command.
Expand All @@ -27,11 +27,11 @@ Plus curated MCP servers and the scaffolding CLI that sets it all up in one comm
| 4 | [`.mcp.json`](./.mcp.json) + [`mcp/`](./mcp/) | MCP servers shipped via project-scoped config |
| 5 | [`.claude/skills/`](./.claude/skills/) | Two skills: `wordpress-scaffold` (greenfield plugin) and `wordpress-feature` (feature + maintenance). Each writes plan artifacts before generating code. Plus the [WordPress/agent-skills](https://github.com/WordPress/agent-skills) library pulled fresh on scaffold |
| 6 | [`.claude/plans/`](./.claude/plans/) | Durable cross-session memory — `constitution.md` (strict for security, default for dependencies), per-feature `spec.md` / `plan.md` / `progress.md`, archived plans |
| 7 | [`.claude/agents/`](./.claude/agents/) | Two read-only sub-agents: `plan-reviewer` (audits spec/plan) and `security-reviewer` (audits code). Block / REST work is covered by skills. |
| 7 | [`.claude/agents/`](./.claude/agents/) | Three read-only sub-agents: `plan-reviewer` (audits spec/plan), `security-reviewer` (audits code), `playground-verifier` (boots the plugin in real WordPress and checks it actually runs). Block / REST work is covered by skills. |
| 8 | [`.claude/hooks/`](./.claude/hooks/) | Pre-commit gate (blocking), post-edit lint, `UserPromptSubmit` plan injection, `Stop` progress timestamp, `SessionStart` orientation banner |
| 9 | [`.claude/commands/`](./.claude/commands/) | Slash commands wrapping common operations — `/plan-freeze`, `/audit-plan`, `/ship-feature` |
| 10 | [`scripts/`](./scripts/) | `quality.sh` — single source of truth for "what is quality"; `open-plan-pr.sh` — plan-freeze automation |
| 11 | [`tests/`](./tests/) | Bash test suite for the kit's own hooks. Integrated into `quality.sh`. |
| 11 | [`tests/`](./tests/) | Bash test suite for the kit's own hooks and sub-agent contracts. Integrated into `quality.sh`. |
| 12 | [`docs/`](./docs/) | Nine-chapter walkthrough from zero to a fully configured harness |

## Quick start
Expand Down
27 changes: 24 additions & 3 deletions docs/06-sub-agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ In the AI Fluency Framework, sub-agents make **Discernimento** concrete: indepen

## The kit's sub-agents

The kit ships two — both read-only audits. The bar for shipping more is high: a sub-agent has to do something a skill can't.
The kit ships three — all read-only. The bar for shipping more is high: a sub-agent has to do something a skill can't.

### `plan-reviewer`

Expand All @@ -28,9 +28,30 @@ Use it: before opening a feature PR, before tagging a release, after touching in

The read-only restriction is the whole point of both. A skill could *describe* the audit; only a sub-agent with a restricted tool list can *enforce* that the audit doesn't turn into a drive-by refactor. That structural guarantee is what makes it worth the extra hop.

### Why not block-builder / rest-builder?
### `playground-verifier`

Tempting, but skills cover that ground more cheaply. A sub-agent without tool restrictions is just a skill with extra ceremony — same context isolation cost, same prompt-as-source-of-truth, no enforcement. Block and REST work is handled by the `wp-block-development` and `wp-rest-api` skills the kit pulls in from [WordPress/agent-skills](https://github.com/WordPress/agent-skills).
Read-only, plus the `wp-playground` MCP tools. Boots the plugin in an ephemeral WordPress instance **at the versions the plugin's own header declares**, mounts the working tree, then activates it, checks the error log, confirms every `register_rest_route` call actually resolves at runtime, deactivates, uninstalls, and verifies no options are left behind. Reports with the same `severity · evidence · suggested fix` shape as the other two, and always tears the instance down.

Use it: before opening a feature PR (in parallel with `security-reviewer`), and immediately after a scaffold.

This one exists because of a real failure. Kit v1.0.2 fixed a scaffold that fataled the moment you activated it — the Composer autoload was PSR-4 while the files used WordPress's `class-{name}.php` convention. It passed phpcs, passed the security review, and shipped in two releases, because **nothing in `quality.sh` boots WordPress**. Static analysis proves the code is well-formed. Only running it proves it works.

### Why not block-builder / rest-builder / planner / coder / tester?

Tempting, and it's the first structure most people reach for — one agent per phase of the development process. But a sub-agent without tool restrictions is just a skill with extra ceremony: same context isolation cost, same prompt-as-source-of-truth, no enforcement.

Run each candidate through the three reasons at the top of this page and most of them collapse:

| Candidate | Verdict |
|---|---|
| Planner | Already a skill (`wordpress-feature` writes spec + plan) — and `plan-reviewer` audits the output |
| Coder / block-builder / rest-builder | Needs the main thread's context and hands back a diff, not a conclusion. Covered by the `wp-block-development` and `wp-rest-api` skills the kit pulls from [WordPress/agent-skills](https://github.com/WordPress/agent-skills) |
| Test-writer | Needs `Edit`. No tool restriction to enforce, so: skill |
| Test-runner / verifier | **Clears the bar** — needs MCP tools the main agent shouldn't hold, dumps thousands of log lines, must not be able to "fix" what it finds. This is `playground-verifier` |
| Reviewer | Clears the bar twice over — `plan-reviewer`, `security-reviewer` |
| Releaser | Deterministic. That's a script and a hook, not a judgement call |

The pattern: delegate work whose **output is a conclusion**, not work whose output is a diff. A conclusion survives the trip back across the context boundary; a diff written by an agent you couldn't watch does not.

## Tool restriction is the lever

Expand Down
Loading
Loading