Repository navigation
Conversation
Adds `test/ai-development/evals`, a standalone workspace that runs coding agents (Claude and Codex) against isolated temporary checkouts and grades the result, so we can measure whether the skills in this repository are discovered and followed. Promptfoo drives the prompt x provider x test matrix. Lifecycle hooks in `lib/workspace-extension.mjs` create a clean Git workspace from HEAD before each row and delete it afterwards; shared provider and assertion defaults live in `lib/`, and each suite under `specs/` supplies its own cases. The workspace also regenerates `.claude/skills` from `.agents/skills`. A real checkout gets that from `npm run agents:setup` via postinstall, but the generated directory is ignored by Git and the workspace never runs npm — so without it the repository's skills would sit there as files nothing announces, and every suite would be measuring a broken environment rather than the guidance under test. The first suite covers testing-skill routing. `npm run test:grader` unit-tests any suite's deterministic grader without spending agent tokens. Run the evals with `npm run test:agent-evals`.
The Claude provider listed neither the skills to enable nor the tool that invokes them, so the repository's skills were never offered to the model. Runs looked like an agent ignoring its guidance when it had never been shown any. Two settings are needed and they do different jobs. `skills` lists skills to the model; omitting it is not "skills off" but no SDK configuration at all. `custom_allowed_tools` replaces the allowed tool list outright rather than extending it, so `Skill` has to appear there for the listed skills to be invocable. With both in place the agent invokes the testing skill and follows it to the e2e reference, leaving the Jest and PHPUnit references unread.
Two corrections to the testing-skill-routing suite, both found by running it. The skill assertion matched a shell command reading `SKILL.md`, so it could only pass if the agent opened the file by hand. Claude invokes the skill natively instead, which left the assertion failing on a run where the skill was used correctly. Promptfoo's `skill-used` matches the invocation itself. The rubric also required asserting serialized content without `expect.poll()`. Nothing in the testing skill says that, and the contributor documentation teaches the opposite: `docs/contributors/code/e2e/overusing-snapshots.md` presents `expect.poll( editor.getBlocks )` as the recommended pattern, and the suite uses `expect.poll( editor.getEditedPostContent )` 164 times across 14 spec files. Playwright only auto-retries locator assertions, so awaiting the getter first collapses it to a one-shot comparison — the weaker form of the two. The clause failed correct work, so it is gone. Skill invocation is not deterministic: across three completed runs on the fixed configuration the agent invoked the skill twice. Single runs cannot support conclusions about this suite.
The Codex provider had never been run. Its binary is not on a normal non-interactive PATH, so the documented command failed for anyone who tried it, and an untested provider in the matrix reads as a broken harness rather than an unfinished one. Removing it leaves the plain command working with no provider filter. The agent-rubric grader moves to Claude as well, since it was configured to grade through Codex. The shared layout still expects more than one agent — providers live in their own file and assertions match Promptfoo's provider-independent command trajectory steps — so a second agent can be added once someone has actually run it.
The task asked for a test proving a typed `&` stays `&` rather than becoming
`&` in serialized content. That is not how the editor behaves, and should
not be: `escapeAmpersand` in `@wordpress/escape-html` escapes a bare ampersand,
so correct output is `<p>&</p>`. The rubric demanded an assertion that
would fail against correct code, and the agent that noticed and asserted on the
rendered DOM instead was marked wrong for being right.
Centring a paragraph has no such ambiguity. The rubric names the attribute the
editor actually writes — `{"style":{"typography":{"textAlign":"center"}}}` —
and calls out the legacy `{"align":"center"}` form, which is only parsed as
input and is easy to copy from the block fixtures.
Both runs now pass the rubric, leaving skill invocation as the only varying
signal.
The package sat at `test/ai-development/evals`, one level below the `test/*` workspace glob, so a root `npm install` skipped it and it carried the only nested `package-lock.json` in the repository. Anyone running the evals had to discover a second install step first. Moving it to `test/ai-development` matches the other test workspaces, resolves it through the root lockfile, and lets the root script address it by name. The cost is that promptfoo's dependencies now install for every contributor — about 640 additional entries in the root lockfile. They are devDependencies, so nothing reaches a published package or plugin build.
The task moved from ampersand encoding to paragraph centre alignment, but the case kept its old description, which is the label shown in results.
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
|
Size Change: 0 B Total Size: 7.9 MB |
|
Flaky tests detected in 2ae3275. 🔍 Workflow run URL: https://github.com/WordPress/gutenberg/actions/runs/32777259507 As a user I want to be able to create a navigation overlay for a specific navigation block in
|
The sandbox options claimed more than they delivered. A probe run showed the agent could read anything on disk — including this repository's checkout, where the eval's own assertions and rubrics live. Removing the evaluation directory from the workspace copy does not help when the original is still readable. `allowRead` only widens access, and the `allowManaged*Only` locks bind solely when passed as managedSettings, where they are filtered through a restrictive-key allowlist; as plain sandbox options they did nothing. So the source root is now denied explicitly, which a repeat of the probe confirms: reading a spec file returns "Operation not permitted" rather than its contents. `failIfUnavailable` is dropped as well — it already defaults to true whenever the sandbox is enabled. What remains is verified in force: network egress fails, writes stay inside the workspace, and the suite still passes.
Three settings were being applied through our own mechanisms rather than the documented ones. The per-test timeout lived in an environment variable inside the npm script, so it was invisible to anyone reading a spec, could not vary per suite, and was lost when running `promptfoo eval` directly. `evaluateOptions.timeoutMs` does the same job in the spec, verified by watching a deliberately short value abort a run with no environment variable set. `options.bustCache` only suppressed cache reads while still writing entries. `evaluateOptions.cache` covers both: `initializeAgenticCache` returns early when caching is disabled, so the per-test option is redundant. Tracing keeps its `$ref`. Promptfoo resolves JSON schema references through `@apidevtools/json-schema-ref-parser` and documents them for reusable blocks, so this is its own sharing mechanism rather than something incidental.
The extension reproduced what `tools/agents/setup-skills.mjs` does rather than calling it, so the workspace would silently stop matching a real checkout if that script ever did more than copy a directory. It exports `setupSkills` and takes the repository root as an argument, so the workspace can be passed directly; only its CLI wrapper hardcodes the root. A missing source directory now fails instead of returning quietly. A workspace without the generated view still runs, and every suite would report an agent ignoring guidance it was never offered — the most misleading result this harness can produce.
The repository writes American English; the task prompt and rubric had crept into British forms.
Tracing, extensions, providers, defaultTest, evaluateOptions and outputPath are the same for every suite, but each spec declared its own. A second suite could quietly run with a different timeout, repeat count or cache setting than the first, and the difference would not be visible from either file. They now live in `lib/promptfooconfig.yaml` and are pulled in by reference, so a spec carries only what is genuinely its own: what it asks for and what it asserts. Promptfoo resolves JSON schema references, and `file://` paths inside a referenced value still resolve relative to the spec, so the provider and defaultTest files load as before. Results are written to one file per run rather than per suite. Promptfoo's own store is the history that `npm run view` reads; this is a convenience dump.
Promptfoo accepts `.js` configs alongside `.yaml`, and this repository is JavaScript nearly everywhere. More usefully, JavaScript composes: a spec spreads `lib/base.js` and adds only what is its own, where the YAML version repeated a reference per shared key and could not express partial overrides at all. The `file://` indirection for providers and default test options goes too — those are now plain imports.
The YAML carried a `yaml-language-server` schema comment, which editors used to
check the config as it was written. Moving to JavaScript dropped that. Promptfoo
exports its config types, so a JSDoc annotation restores it, matching how the
repository types other configuration — `packages/stylelint-config/index.js` does
the same against stylelint's.
It earns its place immediately: `acceptFormats: [ 'json' ]` widened to
`string[]` where the schema wants `('json' | 'protobuf')[]`, which nothing would
have caught before.
Promptfoo compares prompts and providers, but has no notion of the tree an agent works in. This adds that axis, so a change to the repository's agent guidance can be measured rather than assumed. Every workspace is built from a Git ref. A plain run measures the branch you are on, unchanged from before. A comparison run puts two refs in one table, one row per case per ref, graded by the identical assertions — so the difference between the rows is what the change bought. A skill the model already agrees with scores the same on both sides, which is worth knowing before maintaining it. Refs resolve to a commit before any model call, so an unknown ref fails immediately and a moving branch cannot shift underneath a run.
2ae3275 to
6ba2a4d
Compare
12711be to
ac47589
Compare
Important
I haven't reviewed this yet - an idea that I don't want to forget about
Follow up to #80812, which it is based on. Review that first.
What?
Lets an eval run against two Git refs in one table, so a change to the repository's agent guidance can be measured against the commit before it.
Why?
#80812 measures one state of the repository: whatever you have checked out. That tells you whether guidance was followed, but not whether it helped. A skill the model already agrees with scores well while contributing nothing, and a single run cannot tell those apart.
Raised on #81400 — "it's quite hard to understand their impact".
How?
Promptfoo compares prompts and providers, but has no notion of the tree an agent works in, so this supplies that axis.
baseRefvar, and the workspace is built from it. Refs resolve to a commit before any model call, so an unknown ref fails immediately and a moving branch cannot shift underneath a run.withRefs()fromlib/with-refs.js, which expands them across the refs inEVAL_REFS— one row per case per ref.Testing Instructions
npm run test:agent-evals -- --config specs/testing-skill-routing/test-skill-routing.test.js— confirm it still produces a single row, unchanged from Add promptfoo for testing agent development setup and skills #80812.EVAL_REFS=HEAD,trunk npm --workspace @wordpress/agent-skill-evals run eval:compare -- --config specs/testing-skill-routing/test-skill-routing.test.js… @ HEADand… @ trunk, with abaseRefcolumn and the same metrics on both.EVAL_REFS=HEAD,nope npm --workspace @wordpress/agent-skill-evals run eval:compare -- --config …— confirm it fails withRef not found: nopebefore any model call.Each run makes real model calls against your own Claude quota. Results vary between runs; the agent does not always invoke the skill, so a differing result is expected.
Testing Instructions for Keyboard
Not applicable; this PR does not change the user interface.
Screenshots or screencast
Not applicable.
Use of AI Tools
Claude Code (Opus) was used to implement and document this, and Codex (GPT-5) earlier in the series. All output was reviewed and verified by running the harness.