Repository navigation
Conversation
Adds `test/ai-development/evals`, a standalone workspace that runs coding agents (Claude and Codex) against isolated temporary checkouts and grades the result, so we can measure whether the skills in this repository are discovered and followed. Promptfoo drives the prompt x provider x test matrix. Lifecycle hooks in `lib/workspace-extension.mjs` create a clean Git workspace from HEAD before each row and delete it afterwards; shared provider and assertion defaults live in `lib/`, and each suite under `specs/` supplies its own cases. The workspace also regenerates `.claude/skills` from `.agents/skills`. A real checkout gets that from `npm run agents:setup` via postinstall, but the generated directory is ignored by Git and the workspace never runs npm — so without it the repository's skills would sit there as files nothing announces, and every suite would be measuring a broken environment rather than the guidance under test. The first suite covers testing-skill routing. `npm run test:grader` unit-tests any suite's deterministic grader without spending agent tokens. Run the evals with `npm run test:agent-evals`.
The Claude provider listed neither the skills to enable nor the tool that invokes them, so the repository's skills were never offered to the model. Runs looked like an agent ignoring its guidance when it had never been shown any. Two settings are needed and they do different jobs. `skills` lists skills to the model; omitting it is not "skills off" but no SDK configuration at all. `custom_allowed_tools` replaces the allowed tool list outright rather than extending it, so `Skill` has to appear there for the listed skills to be invocable. With both in place the agent invokes the testing skill and follows it to the e2e reference, leaving the Jest and PHPUnit references unread.
Two corrections to the testing-skill-routing suite, both found by running it. The skill assertion matched a shell command reading `SKILL.md`, so it could only pass if the agent opened the file by hand. Claude invokes the skill natively instead, which left the assertion failing on a run where the skill was used correctly. Promptfoo's `skill-used` matches the invocation itself. The rubric also required asserting serialized content without `expect.poll()`. Nothing in the testing skill says that, and the contributor documentation teaches the opposite: `docs/contributors/code/e2e/overusing-snapshots.md` presents `expect.poll( editor.getBlocks )` as the recommended pattern, and the suite uses `expect.poll( editor.getEditedPostContent )` 164 times across 14 spec files. Playwright only auto-retries locator assertions, so awaiting the getter first collapses it to a one-shot comparison — the weaker form of the two. The clause failed correct work, so it is gone. Skill invocation is not deterministic: across three completed runs on the fixed configuration the agent invoked the skill twice. Single runs cannot support conclusions about this suite.
The Codex provider had never been run. Its binary is not on a normal non-interactive PATH, so the documented command failed for anyone who tried it, and an untested provider in the matrix reads as a broken harness rather than an unfinished one. Removing it leaves the plain command working with no provider filter. The agent-rubric grader moves to Claude as well, since it was configured to grade through Codex. The shared layout still expects more than one agent — providers live in their own file and assertions match Promptfoo's provider-independent command trajectory steps — so a second agent can be added once someone has actually run it.
The task asked for a test proving a typed `&` stays `&` rather than becoming
`&` in serialized content. That is not how the editor behaves, and should
not be: `escapeAmpersand` in `@wordpress/escape-html` escapes a bare ampersand,
so correct output is `<p>&</p>`. The rubric demanded an assertion that
would fail against correct code, and the agent that noticed and asserted on the
rendered DOM instead was marked wrong for being right.
Centring a paragraph has no such ambiguity. The rubric names the attribute the
editor actually writes — `{"style":{"typography":{"textAlign":"center"}}}` —
and calls out the legacy `{"align":"center"}` form, which is only parsed as
input and is easy to copy from the block fixtures.
Both runs now pass the rubric, leaving skill invocation as the only varying
signal.
The package sat at `test/ai-development/evals`, one level below the `test/*` workspace glob, so a root `npm install` skipped it and it carried the only nested `package-lock.json` in the repository. Anyone running the evals had to discover a second install step first. Moving it to `test/ai-development` matches the other test workspaces, resolves it through the root lockfile, and lets the root script address it by name. The cost is that promptfoo's dependencies now install for every contributor — about 640 additional entries in the root lockfile. They are devDependencies, so nothing reaches a published package or plugin build.
The task moved from ampersand encoding to paragraph centre alignment, but the case kept its old description, which is the label shown in results.
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
|
Size Change: 0 B Total Size: 7.9 MB |
jeryj
marked this pull request as draft
August 25, 2026 13:41
The sandbox options claimed more than they delivered. A probe run showed the agent could read anything on disk — including this repository's checkout, where the eval's own assertions and rubrics live. Removing the evaluation directory from the workspace copy does not help when the original is still readable. `allowRead` only widens access, and the `allowManaged*Only` locks bind solely when passed as managedSettings, where they are filtered through a restrictive-key allowlist; as plain sandbox options they did nothing. So the source root is now denied explicitly, which a repeat of the probe confirms: reading a spec file returns "Operation not permitted" rather than its contents. `failIfUnavailable` is dropped as well — it already defaults to true whenever the sandbox is enabled. What remains is verified in force: network egress fails, writes stay inside the workspace, and the suite still passes.
Three settings were being applied through our own mechanisms rather than the documented ones. The per-test timeout lived in an environment variable inside the npm script, so it was invisible to anyone reading a spec, could not vary per suite, and was lost when running `promptfoo eval` directly. `evaluateOptions.timeoutMs` does the same job in the spec, verified by watching a deliberately short value abort a run with no environment variable set. `options.bustCache` only suppressed cache reads while still writing entries. `evaluateOptions.cache` covers both: `initializeAgenticCache` returns early when caching is disabled, so the per-test option is redundant. Tracing keeps its `$ref`. Promptfoo resolves JSON schema references through `@apidevtools/json-schema-ref-parser` and documents them for reusable blocks, so this is its own sharing mechanism rather than something incidental.
The extension reproduced what `tools/agents/setup-skills.mjs` does rather than calling it, so the workspace would silently stop matching a real checkout if that script ever did more than copy a directory. It exports `setupSkills` and takes the repository root as an argument, so the workspace can be passed directly; only its CLI wrapper hardcodes the root. A missing source directory now fails instead of returning quietly. A workspace without the generated view still runs, and every suite would report an agent ignoring guidance it was never offered — the most misleading result this harness can produce.
The repository writes American English; the task prompt and rubric had crept into British forms.
Tracing, extensions, providers, defaultTest, evaluateOptions and outputPath are the same for every suite, but each spec declared its own. A second suite could quietly run with a different timeout, repeat count or cache setting than the first, and the difference would not be visible from either file. They now live in `lib/promptfooconfig.yaml` and are pulled in by reference, so a spec carries only what is genuinely its own: what it asks for and what it asserts. Promptfoo resolves JSON schema references, and `file://` paths inside a referenced value still resolve relative to the spec, so the provider and defaultTest files load as before. Results are written to one file per run rather than per suite. Promptfoo's own store is the history that `npm run view` reads; this is a convenience dump.
Promptfoo accepts `.js` configs alongside `.yaml`, and this repository is JavaScript nearly everywhere. More usefully, JavaScript composes: a spec spreads `lib/base.js` and adds only what is its own, where the YAML version repeated a reference per shared key and could not express partial overrides at all. The `file://` indirection for providers and default test options goes too — those are now plain imports.
The YAML carried a `yaml-language-server` schema comment, which editors used to
check the config as it was written. Moving to JavaScript dropped that. Promptfoo
exports its config types, so a JSDoc annotation restores it, matching how the
repository types other configuration — `packages/stylelint-config/index.js` does
the same against stylelint's.
It earns its place immediately: `acceptFormats: [ 'json' ]` widened to
`string[]` where the schema wants `('json' | 'protobuf')[]`, which nothing would
have caught before.
Claude is the only agent in the harness. Running two over the same suite is the point of the matrix: guidance that works for one agent and not the other is worth knowing about, and is invisible with a single provider. No suite changes are needed. Assertions match Promptfoo's provider-independent command trajectory steps rather than one agent's tool names, and the shared tracing config already declares Claude's `Bash` alongside Codex's `exec_command`. The `codex` binary is not on a normal PATH, so the README says where it comes from and a failure reads as a missing binary rather than a broken eval.
jeryj
force-pushed
the
add/promptfoo-codex
branch
from
August 25, 2026 18:39
bc763fb to
9c718dc
Compare
ciampo
force-pushed
the
add/promptfoo
branch
8 times, most recently
from
August 31, 2026 18:45
ac47589 to
059432e
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Important
I haven't reviewed this yet - an idea that I don't want to forget about
Follow up to #80812, which it is based on. Review that first.
What?
Adds Codex as a second coding agent in the eval matrix, alongside Claude.
Why?
#80812 ships Claude only. Running two agents over the same suite is the point of the matrix: guidance that works for one and not the other is worth knowing about, and is invisible with a single provider.
Running both agents over the same suite is the point of the matrix: guidance that works for one agent and not the other is worth knowing about, and is invisible with a single provider.
How?
openai:codex-sdkprovider, sandboxed to the workspace with network access and web search disabled.codexbinary requirement in the README, so a failure is recognisable as a missing binary rather than a broken eval.No suite changes are needed. Assertions already match Promptfoo's provider-independent command trajectory steps rather than one agent's tool names, and the shared tracing config already declares Claude's
Bashalongside Codex'sexec_command.Testing Instructions
codexis on yourPATH(which codex). If it is not, this PR cannot be tested — that is the constraint being documented.npm run test:agent-evals -- --config specs/testing-skill-routing/test-skill-routing.test.jsnpm --workspace @wordpress/agent-skill-evals run viewclaudeandcodex, with the same metrics evaluated for both.npm run test:agent-evals -- --config … --filter-providers codex— confirm a single-agent run works.Each run makes real model calls against your own quota. Results vary between runs; the agent does not always invoke the skill, so a differing result is expected rather than a broken PR.
Testing Instructions for Keyboard
Not applicable; this PR does not change the user interface.
Screenshots or screencast
Not applicable.
Use of AI Tools
Claude Code (Opus) was used to implement and document this, and Codex (GPT-5) earlier in the series. All output was reviewed and verified by running the harness.