Repository navigation
Conversation
Verifies that an agent, given a scenario's task prompt, consults the right guidance: reads the skill files and docs it should, doesn't read the ones it shouldn't, in the right order, and reaches for the right kind of command — with the evidence coming from the agent's transcript. - eval/harness/run.mjs: static sanity checks (skill frontmatter, scenario files parse) plus a live mode that runs scenario queries through a real agent session headlessly, normalizes the transcript into read/write/command events, and evaluates declarative assertions with pass rates (--scenario, --model, --agent, --repeat). Needs no running environment and grants no write permissions — a denied edit attempt is still evidence, and the checkout is never modified. - eval/scenarios/: two scenarios distilled from transcript-validated manual runs against the testing skill, plus a negative control that guards against over-triggering (trivial tasks must not read skills or contributor guides). - docs/contributors/code/agents-and-skills.md: adding a skill now includes adding an eval scenario. Verified live: testing-write-e2e passes all six assertions end to end; the negative control passes all three. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Relocate eval/ to test/ai-development/ — test infrastructure lives under test/ — flattening the harness/ subdirectory. Scenario and artifact paths now resolve relative to the script itself. - Format run.mjs with the repo formatter and allow console output for this CLI script. The initial commit never met the formatter because lint-staged's patterns do not cover .mjs files; CI's lint job caught it. - Update path references in the scenarios README and the agents-and-skills.md checklist. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Alternative to the standalone runner in #80703. The core survives unchanged in test/ai-development/agent.mjs: the claude adapter, transcript normalization into read/write/command events, and the declarative assertion evaluation. @playwright/test (already installed, runner only — no browsers) replaces the bespoke runner: repeats via --repeat-each, filtering via -g, standard reporting and exit codes. The static frontmatter checks are dropped. Run with `npm run test:ai-development`; select a model with AI_EVAL_MODEL. Verified: the negative-control scenario passes end to end through the new pipeline, and `npm run test:unit -- --listTests` confirms Jest never sweeps these tests up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
|
Size Change: 0 B Total Size: 7.75 MB |
|
Flaky tests detected in 1e84c8f. 🔍 Workflow run URL: https://github.com/WordPress/gutenberg/actions/runs/30299109106
|
Following review: delete the JSON scenario files and the custom matcher layer; tests are now hand-written specs using stock expect. The adapter reports repo-relative file paths, so reads assert with exact matches ( expect( transcript.reads ).toContain( 'skills/testing/SKILL.md' ) ), directory exclusions are a filter, and ordering is a comparison of first-event indexes (Infinity when absent). Standard timeouts: 10-minute agent sessions in agent.mjs, 12-minute test ceiling in the config — the session timeout is the real backstop since spawnSync blocks the worker. Verified: negative control passes live through the final API; lint and formatting clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
| claude: { | ||
| command: 'claude', | ||
| buildArgs( query, model ) { | ||
| const args = [ |
There was a problem hiding this comment.
This doesn't restrict writes, so technically Claude could modify the code, even though the README says:
The agent is granted no write permissions
| expect | ||
| .soft( | ||
| transcript.reads.filter( ( r ) => | ||
| r.startsWith( 'docs/contributors/code/e2e/' ) |
There was a problem hiding this comment.
Should this be:
| r.startsWith( 'docs/contributors/code/e2e/' ) | |
| r.startsWith( 'docs/contributors/code/' ) |
so that we notice reads from all these files?
| evidence.push( | ||
| fileEvidence( 'write', input.file_path ) | ||
| ); | ||
| } else if ( block.name === 'Bash' && input.command ) { |
There was a problem hiding this comment.
Should we also push fileEvidence in this branch?
| ); | ||
| } else if ( block.name === 'Bash' && input.command ) { | ||
| evidence.push( commandEvidence( input.command ) ); | ||
| } |
There was a problem hiding this comment.
Do we also need a branch for Grep?
| // The guide must be consulted before the first attempted spec edit. | ||
| expect | ||
| .soft( | ||
| transcript.firstRead( 'docs/contributors/code/e2e/README.md' ) | ||
| ) | ||
| .toBeLessThan( transcript.firstWrite() ); |
There was a problem hiding this comment.
Could we also record shell-based edits as writes here? This check currently only sees writes reported as file-access events. For example, when an agent uses apply_patch, we record the command but firstWrite() stays Infinity.
That means an agent could edit first and read the guide afterwards, and the assertion would still pass. Would it make sense to either record shell edits as writes or use an ordering check that catches them too?
What?
Add Playwright tests for evaluating AI success criteria, so we can iterate and improve our agent set-up and skills with more confidence.
Why?
As we grow our AGENTS.md and skills in Gutenberg, we'll want to make sure we're not accidentally causing regressions. We also want to ensure our skills perform what we want them to, and eventually add things to evaluate token cost and context usage.
How?
test/ai-development/agent.mjs— the single helper:runAgent( query )spawns a headless agent session and normalizes its transcript into ordered evidence records (file reads, writes, and commands executed). Normalized via Adapters mapped by the agents (Claude and Codex currently supported).test/ai-development/specs/— specs written directly with Playwright's built-in matchers to evaluate the agent's behavior.test/ai-development/artifacts/.docs/contributors/code/agents-and-skills.md: the add-a-skill checklist points at these tests.Testing Instructions
Requires the
claudeCLI signed in. Each step spawns real agent sessions (minutes and real tokens).npm run test:ai-development -- -g "trivial task"(~1 minute) — confirm the terminal reports1 passed.npm run test:ai-development(~5–15 minutes) — confirm the terminal lists all three tests with a pass/fail each; a failure prints the expected value against the full list of files the agent actually read.git status— confirm a clean tree (transcripts land only in the gitignoredtest/ai-development/artifacts/).Common options:
-g "<test name>",--repeat-each=3(compliance is probabilistic),AI_EVAL_MODEL=haikuto target another model,AI_EVAL_AGENT=codexto target Codex.Use of AI Tools
Written with Claude Code (Claude Fable 5) in an interactive session, directed and reviewed by the PR author.
🤖 Generated with Claude Code