Skip to content

Agent evals: Compare guidance changes across two Git refs - #82000

Draft
jeryj wants to merge 15 commits into
trunkfrom
add/promptfoo-comparison
Draft

jeryj wants to merge 15 commits into
trunkfrom
add/promptfoo-comparison

Conversation

@jeryj

@jeryj jeryj commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Important

I haven't reviewed this yet - an idea that I don't want to forget about

Follow up to #80812, which it is based on. Review that first.

What?

Lets an eval run against two Git refs in one table, so a change to the repository's agent guidance can be measured against the commit before it.

Why?

#80812 measures one state of the repository: whatever you have checked out. That tells you whether guidance was followed, but not whether it helped. A skill the model already agrees with scores well while contributing nothing, and a single run cannot tell those apart.

Raised on #81400 — "it's quite hard to understand their impact".

How?

Promptfoo compares prompts and providers, but has no notion of the tree an agent works in, so this supplies that axis.

  • Each row carries a baseRef var, and the workspace is built from it. Refs resolve to a commit before any model call, so an unknown ref fails immediately and a moving branch cannot shift underneath a run.
  • A suite wraps its cases in withRefs() from lib/with-refs.js, which expands them across the refs in EVAL_REFS — one row per case per ref.
  • Both refs are graded by the identical assertions, so the difference between rows is what the change bought.

Testing Instructions

  1. npm run test:agent-evals -- --config specs/testing-skill-routing/test-skill-routing.test.js — confirm it still produces a single row, unchanged from Add promptfoo for testing agent development setup and skills #80812.
  2. EVAL_REFS=HEAD,trunk npm --workspace @wordpress/agent-skill-evals run eval:compare -- --config specs/testing-skill-routing/test-skill-routing.test.js
  3. Confirm the results table has one row per ref, each labelled … @ HEAD and … @ trunk, with a baseRef column and the same metrics on both.
  4. EVAL_REFS=HEAD,nope npm --workspace @wordpress/agent-skill-evals run eval:compare -- --config … — confirm it fails with Ref not found: nope before any model call.

Each run makes real model calls against your own Claude quota. Results vary between runs; the agent does not always invoke the skill, so a differing result is expected.

Testing Instructions for Keyboard

Not applicable; this PR does not change the user interface.

Screenshots or screencast

Not applicable.

Use of AI Tools

Claude Code (Opus) was used to implement and document this, and Codex (GPT-5) earlier in the series. All output was reviewed and verified by running the harness.

jeryj added 7 commits August 24, 2026 10:51
Adds `test/ai-development/evals`, a standalone workspace that runs coding
agents (Claude and Codex) against isolated temporary checkouts and grades
the result, so we can measure whether the skills in this repository are
discovered and followed.

Promptfoo drives the prompt x provider x test matrix. Lifecycle hooks in
`lib/workspace-extension.mjs` create a clean Git workspace from HEAD before
each row and delete it afterwards; shared provider and assertion defaults
live in `lib/`, and each suite under `specs/` supplies its own cases.

The workspace also regenerates `.claude/skills` from `.agents/skills`. A real
checkout gets that from `npm run agents:setup` via postinstall, but the
generated directory is ignored by Git and the workspace never runs npm — so
without it the repository's skills would sit there as files nothing announces,
and every suite would be measuring a broken environment rather than the
guidance under test.

The first suite covers testing-skill routing. `npm run test:grader`
unit-tests any suite's deterministic grader without spending agent tokens.

Run the evals with `npm run test:agent-evals`.
The Claude provider listed neither the skills to enable nor the tool that
invokes them, so the repository's skills were never offered to the model. Runs
looked like an agent ignoring its guidance when it had never been shown any.

Two settings are needed and they do different jobs. `skills` lists skills to
the model; omitting it is not "skills off" but no SDK configuration at all.
`custom_allowed_tools` replaces the allowed tool list outright rather than
extending it, so `Skill` has to appear there for the listed skills to be
invocable.

With both in place the agent invokes the testing skill and follows it to the
e2e reference, leaving the Jest and PHPUnit references unread.
Two corrections to the testing-skill-routing suite, both found by running it.

The skill assertion matched a shell command reading `SKILL.md`, so it could
only pass if the agent opened the file by hand. Claude invokes the skill
natively instead, which left the assertion failing on a run where the skill was
used correctly. Promptfoo's `skill-used` matches the invocation itself.

The rubric also required asserting serialized content without `expect.poll()`.
Nothing in the testing skill says that, and the contributor documentation
teaches the opposite: `docs/contributors/code/e2e/overusing-snapshots.md`
presents `expect.poll( editor.getBlocks )` as the recommended pattern, and the
suite uses `expect.poll( editor.getEditedPostContent )` 164 times across 14
spec files. Playwright only auto-retries locator assertions, so awaiting the
getter first collapses it to a one-shot comparison — the weaker form of the
two. The clause failed correct work, so it is gone.

Skill invocation is not deterministic: across three completed runs on the fixed
configuration the agent invoked the skill twice. Single runs cannot support
conclusions about this suite.
The Codex provider had never been run. Its binary is not on a normal
non-interactive PATH, so the documented command failed for anyone who tried it,
and an untested provider in the matrix reads as a broken harness rather than an
unfinished one.

Removing it leaves the plain command working with no provider filter. The
agent-rubric grader moves to Claude as well, since it was configured to grade
through Codex.

The shared layout still expects more than one agent — providers live in their
own file and assertions match Promptfoo's provider-independent command
trajectory steps — so a second agent can be added once someone has actually
run it.
The task asked for a test proving a typed `&` stays `&` rather than becoming
`&` in serialized content. That is not how the editor behaves, and should
not be: `escapeAmpersand` in `@wordpress/escape-html` escapes a bare ampersand,
so correct output is `<p>&amp;</p>`. The rubric demanded an assertion that
would fail against correct code, and the agent that noticed and asserted on the
rendered DOM instead was marked wrong for being right.

Centring a paragraph has no such ambiguity. The rubric names the attribute the
editor actually writes — `{"style":{"typography":{"textAlign":"center"}}}` —
and calls out the legacy `{"align":"center"}` form, which is only parsed as
input and is easy to copy from the block fixtures.

Both runs now pass the rubric, leaving skill invocation as the only varying
signal.
The package sat at `test/ai-development/evals`, one level below the `test/*`
workspace glob, so a root `npm install` skipped it and it carried the only
nested `package-lock.json` in the repository. Anyone running the evals had to
discover a second install step first.

Moving it to `test/ai-development` matches the other test workspaces, resolves
it through the root lockfile, and lets the root script address it by name.

The cost is that promptfoo's dependencies now install for every contributor —
about 640 additional entries in the root lockfile. They are devDependencies, so
nothing reaches a published package or plugin build.
The task moved from ampersand encoding to paragraph centre alignment, but the
case kept its old description, which is the label shown in results.
@github-actions

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: jeryj <jeryj@git.wordpress.org>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@github-actions

Copy link
Copy Markdown

Size Change: 0 B

Total Size: 7.9 MB

compressed-size-action

@github-actions

Copy link
Copy Markdown

Flaky tests detected in 2ae3275.
Some tests passed with failed attempts. The failures may not be related to this commit but are still reported for visibility. See the documentation for more information.

🔍 Workflow run URL: https://github.com/WordPress/gutenberg/actions/runs/32777259507
📝 Reported tests:

As a user I want to be able to create a navigation overlay for a specific navigation block in /test/e2e/specs/site-editor/navigation-overlay-template-part.spec.js, passed after 2 failed attempts.
Error: apiRequestContext.fetch: socket hang up
Call log:
  - → GET http://localhost:8889/wp-json/wp/v2/templates
    - user-agent: Playwright/1.62.1 (x64; ubuntu 24.04) node/20.20 CI/1
    - accept: */*
    - accept-encoding: gzip,deflate,br
    - X-WP-Nonce: de30f1ca08
    - cookie: wordpress_test_cookie=WP%20Cookie%20check; wordpress_logged_in_23778236db82f19306f247e20a353a99=admin%7C1787778580%7CDHlCr3VxqtOU3kTWTXnN83nEHYRPRLLuiM0l77YaKVh%7Cc20312cb8b0d33ae61c02945ba8ad0edb2c51e204c5c5f0b889700ea952a747a; wp-settings-time-1=1787606479; wp-settings-1=editor%3Dtinymce

    at RequestUtils.rest (/home/runner/work/gutenberg/gutenberg/packages/e2e-test-utils-playwright/src/request-utils/rest.ts:112:39)
    at RequestUtils.deleteAllTemplates (/home/runner/work/gutenberg/gutenberg/packages/e2e-test-utils-playwright/src/request-utils/templates.ts:35:31)
    at /home/runner/work/gutenberg/gutenberg/test/e2e/specs/site-editor/navigation-overlay-template-part.spec.js:57:22
TimeoutError: locator.click: Timeout 10000ms exceeded.
Call log:
  - waiting for getByRole('button', { name: 'Create overlay', exact: true })

    at createNavigationOverlay (/home/runner/work/gutenberg/gutenberg/test/e2e/specs/site-editor/navigation-overlay-template-part.spec.js:44:28)
    at /home/runner/work/gutenberg/gutenberg/test/e2e/specs/site-editor/navigation-overlay-template-part.spec.js:75:4
should allow user to add and remove multiple local font files in /test/e2e/specs/admin/font-library.spec.js, passed after 1 failed attempt.
Error: expect(locator).toBeVisible() failed

Locator: getByLabel('Exo 2 Semi-bold Italic')
Expected: visible
Timeout: 5000ms
Error: element(s) not found

Call log:
  - Expect "toBeVisible" with timeout 5000ms
  - waiting for getByLabel('Exo 2 Semi-bold Italic')

    at /home/runner/work/gutenberg/gutenberg/test/e2e/specs/admin/font-library.spec.js:58:6

@jeryj
jeryj marked this pull request as draft August 25, 2026 13:38
@jeryj jeryj self-assigned this Aug 25, 2026
@jeryj jeryj added the [Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests. label Aug 25, 2026
jeryj added 2 commits August 25, 2026 09:05
The sandbox options claimed more than they delivered. A probe run showed the
agent could read anything on disk — including this repository's checkout, where
the eval's own assertions and rubrics live. Removing the evaluation directory
from the workspace copy does not help when the original is still readable.

`allowRead` only widens access, and the `allowManaged*Only` locks bind solely
when passed as managedSettings, where they are filtered through a
restrictive-key allowlist; as plain sandbox options they did nothing. So the
source root is now denied explicitly, which a repeat of the probe confirms:
reading a spec file returns "Operation not permitted" rather than its contents.

`failIfUnavailable` is dropped as well — it already defaults to true whenever
the sandbox is enabled.

What remains is verified in force: network egress fails, writes stay inside the
workspace, and the suite still passes.
Three settings were being applied through our own mechanisms rather than the
documented ones.

The per-test timeout lived in an environment variable inside the npm script, so
it was invisible to anyone reading a spec, could not vary per suite, and was
lost when running `promptfoo eval` directly. `evaluateOptions.timeoutMs` does
the same job in the spec, verified by watching a deliberately short value abort
a run with no environment variable set.

`options.bustCache` only suppressed cache reads while still writing entries.
`evaluateOptions.cache` covers both: `initializeAgenticCache` returns early
when caching is disabled, so the per-test option is redundant.

Tracing keeps its `$ref`. Promptfoo resolves JSON schema references through
`@apidevtools/json-schema-ref-parser` and documents them for reusable blocks,
so this is its own sharing mechanism rather than something incidental.
jeryj added 6 commits August 25, 2026 11:41
The extension reproduced what `tools/agents/setup-skills.mjs` does rather than
calling it, so the workspace would silently stop matching a real checkout if
that script ever did more than copy a directory. It exports `setupSkills` and
takes the repository root as an argument, so the workspace can be passed
directly; only its CLI wrapper hardcodes the root.

A missing source directory now fails instead of returning quietly. A workspace
without the generated view still runs, and every suite would report an agent
ignoring guidance it was never offered — the most misleading result this
harness can produce.
The repository writes American English; the task prompt and rubric had crept
into British forms.
Tracing, extensions, providers, defaultTest, evaluateOptions and outputPath are
the same for every suite, but each spec declared its own. A second suite could
quietly run with a different timeout, repeat count or cache setting than the
first, and the difference would not be visible from either file.

They now live in `lib/promptfooconfig.yaml` and are pulled in by reference, so a
spec carries only what is genuinely its own: what it asks for and what it
asserts. Promptfoo resolves JSON schema references, and `file://` paths inside a
referenced value still resolve relative to the spec, so the provider and
defaultTest files load as before.

Results are written to one file per run rather than per suite. Promptfoo's own
store is the history that `npm run view` reads; this is a convenience dump.
Promptfoo accepts `.js` configs alongside `.yaml`, and this repository is
JavaScript nearly everywhere. More usefully, JavaScript composes: a spec spreads
`lib/base.js` and adds only what is its own, where the YAML version repeated a
reference per shared key and could not express partial overrides at all.

The `file://` indirection for providers and default test options goes too —
those are now plain imports.
The YAML carried a `yaml-language-server` schema comment, which editors used to
check the config as it was written. Moving to JavaScript dropped that. Promptfoo
exports its config types, so a JSDoc annotation restores it, matching how the
repository types other configuration — `packages/stylelint-config/index.js` does
the same against stylelint's.

It earns its place immediately: `acceptFormats: [ 'json' ]` widened to
`string[]` where the schema wants `('json' | 'protobuf')[]`, which nothing would
have caught before.
Promptfoo compares prompts and providers, but has no notion of the tree an
agent works in. This adds that axis, so a change to the repository's agent
guidance can be measured rather than assumed.

Every workspace is built from a Git ref. A plain run measures the branch you
are on, unchanged from before. A comparison run puts two refs in one table, one
row per case per ref, graded by the identical assertions — so the difference
between the rows is what the change bought. A skill the model already agrees
with scores the same on both sides, which is worth knowing before maintaining
it.

Refs resolve to a commit before any model call, so an unknown ref fails
immediately and a moving branch cannot shift underneath a run.
@jeryj
jeryj force-pushed the add/promptfoo-comparison branch from 2ae3275 to 6ba2a4d Compare August 25, 2026 18:41
@ciampo
ciampo force-pushed the add/promptfoo branch 7 times, most recently from 12711be to ac47589 Compare August 31, 2026 18:36
Base automatically changed from add/promptfoo to trunk September 4, 2026 21:21

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

[Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant