Skip to content

Add agent skill evals via Playwright - #80758

Open
jeryj wants to merge 6 commits into
trunkfrom
alt/ai-development-tests
Open

jeryj wants to merge 6 commits into
trunkfrom
alt/ai-development-tests

Conversation

@jeryj

@jeryj jeryj commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

What?

Add Playwright tests for evaluating AI success criteria, so we can iterate and improve our agent set-up and skills with more confidence.

Why?

As we grow our AGENTS.md and skills in Gutenberg, we'll want to make sure we're not accidentally causing regressions. We also want to ensure our skills perform what we want them to, and eventually add things to evaluate token cost and context usage.

How?

  • test/ai-development/agent.mjs — the single helper: runAgent( query ) spawns a headless agent session and normalizes its transcript into ordered evidence records (file reads, writes, and commands executed). Normalized via Adapters mapped by the agents (Claude and Codex currently supported).
  • test/ai-development/specs/ — specs written directly with Playwright's built-in matchers to evaluate the agent's behavior.
  • Each run's raw transcript is saved under the gitignored test/ai-development/artifacts/.
  • docs/contributors/code/agents-and-skills.md: the add-a-skill checklist points at these tests.

Testing Instructions

Requires the claude CLI signed in. Each step spawns real agent sessions (minutes and real tokens).

  1. Run npm run test:ai-development -- -g "trivial task" (~1 minute) — confirm the terminal reports 1 passed.
  2. Run npm run test:ai-development (~5–15 minutes) — confirm the terminal lists all three tests with a pass/fail each; a failure prints the expected value against the full list of files the agent actually read.
  3. Run git status — confirm a clean tree (transcripts land only in the gitignored test/ai-development/artifacts/).

Common options: -g "<test name>", --repeat-each=3 (compliance is probabilistic), AI_EVAL_MODEL=haiku to target another model, AI_EVAL_AGENT=codex to target Codex.

Use of AI Tools

Written with Claude Code (Claude Fable 5) in an interactive session, directed and reviewed by the PR author.

🤖 Generated with Claude Code

jeryj and others added 4 commits July 24, 2026 15:54
Verifies that an agent, given a scenario's task prompt, consults the
right guidance: reads the skill files and docs it should, doesn't read
the ones it shouldn't, in the right order, and reaches for the right
kind of command — with the evidence coming from the agent's transcript.

- eval/harness/run.mjs: static sanity checks (skill frontmatter,
  scenario files parse) plus a live mode that runs scenario queries
  through a real agent session headlessly, normalizes the transcript
  into read/write/command events, and evaluates declarative assertions
  with pass rates (--scenario, --model, --agent, --repeat). Needs no
  running environment and grants no write permissions — a denied edit
  attempt is still evidence, and the checkout is never modified.
- eval/scenarios/: two scenarios distilled from transcript-validated
  manual runs against the testing skill, plus a negative control that
  guards against over-triggering (trivial tasks must not read skills
  or contributor guides).
- docs/contributors/code/agents-and-skills.md: adding a skill now
  includes adding an eval scenario.

Verified live: testing-write-e2e passes all six assertions end to end;
the negative control passes all three.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Relocate eval/ to test/ai-development/ — test infrastructure lives
  under test/ — flattening the harness/ subdirectory. Scenario and
  artifact paths now resolve relative to the script itself.
- Format run.mjs with the repo formatter and allow console output for
  this CLI script. The initial commit never met the formatter because
  lint-staged's patterns do not cover .mjs files; CI's lint job caught
  it.
- Update path references in the scenarios README and the
  agents-and-skills.md checklist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Alternative to the standalone runner in #80703. The core survives
unchanged in test/ai-development/agent.mjs: the claude adapter,
transcript normalization into read/write/command events, and the
declarative assertion evaluation. @playwright/test (already installed,
runner only — no browsers) replaces the bespoke runner: repeats via
--repeat-each, filtering via -g, standard reporting and exit codes.
The static frontmatter checks are dropped.

Run with `npm run test:ai-development`; select a model with
AI_EVAL_MODEL. Verified: the negative-control scenario passes end to
end through the new pipeline, and `npm run test:unit -- --listTests`
confirms Jest never sweeps these tests up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: jeryj <jeryj@git.wordpress.org>
Co-authored-by: scruffian <scruffian@git.wordpress.org>
Co-authored-by: ciampo <mciampini@git.wordpress.org>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@github-actions

Copy link
Copy Markdown

Size Change: 0 B

Total Size: 7.75 MB

compressed-size-action

@github-actions

Copy link
Copy Markdown

Flaky tests detected in 1e84c8f.
Some tests passed with failed attempts. The failures may not be related to this commit but are still reported for visibility. See the documentation for more information.

🔍 Workflow run URL: https://github.com/WordPress/gutenberg/actions/runs/30299109106
📝 Reported issues:

jeryj and others added 2 commits July 27, 2026 16:32
Following review: delete the JSON scenario files and the custom matcher
layer; tests are now hand-written specs using stock expect. The adapter
reports repo-relative file paths, so reads assert with exact matches
( expect( transcript.reads ).toContain( 'skills/testing/SKILL.md' ) ),
directory exclusions are a filter, and ordering is a comparison of
first-event indexes (Infinity when absent). Standard timeouts: 10-minute
agent sessions in agent.mjs, 12-minute test ceiling in the config —
the session timeout is the real backstop since spawnSync blocks the
worker.

Verified: negative control passes live through the final API; lint and
formatting clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@jeryj jeryj added [Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests. [Type] Enhancement A suggestion for improvement. and removed [Type] Enhancement A suggestion for improvement. labels Jul 28, 2026
@jeryj jeryj changed the title Agent skill evals on the Playwright test runner (alternate to #80703) Add agent skill evals via Playwright Jul 28, 2026
claude: {
command: 'claude',
buildArgs( query, model ) {
const args = [

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This doesn't restrict writes, so technically Claude could modify the code, even though the README says:

The agent is granted no write permissions

expect
.soft(
transcript.reads.filter( ( r ) =>
r.startsWith( 'docs/contributors/code/e2e/' )

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should this be:

Suggested change
r.startsWith( 'docs/contributors/code/e2e/' )
r.startsWith( 'docs/contributors/code/' )

so that we notice reads from all these files?

evidence.push(
fileEvidence( 'write', input.file_path )
);
} else if ( block.name === 'Bash' && input.command ) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we also push fileEvidence in this branch?

);
} else if ( block.name === 'Bash' && input.command ) {
evidence.push( commandEvidence( input.command ) );
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we also need a branch for Grep?

Comment on lines +84 to +89
// The guide must be consulted before the first attempted spec edit.
expect
.soft(
transcript.firstRead( 'docs/contributors/code/e2e/README.md' )
)
.toBeLessThan( transcript.firstWrite() );

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we also record shell-based edits as writes here? This check currently only sees writes reported as file-access events. For example, when an agent uses apply_patch, we record the command but firstWrite() stays Infinity.

That means an agent could edit first and read the guide afterwards, and the assertion would still pass. Would it make sense to either record shell edits as writes or use an ordering check that catches them too?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

[Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants