Skip to content

Agent evals: Add Codex as a second agent - #82001

Draft
jeryj wants to merge 15 commits into
trunkfrom
add/promptfoo-codex
Draft

jeryj wants to merge 15 commits into
trunkfrom
add/promptfoo-codex

Conversation

@jeryj

@jeryj jeryj commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Important

I haven't reviewed this yet - an idea that I don't want to forget about

Follow up to #80812, which it is based on. Review that first.

What?

Adds Codex as a second coding agent in the eval matrix, alongside Claude.

Why?

#80812 ships Claude only. Running two agents over the same suite is the point of the matrix: guidance that works for one and not the other is worth knowing about, and is invisible with a single provider.

Running both agents over the same suite is the point of the matrix: guidance that works for one agent and not the other is worth knowing about, and is invisible with a single provider.

How?

  • Restores the openai:codex-sdk provider, sandboxed to the workspace with network access and web search disabled.
  • Documents the codex binary requirement in the README, so a failure is recognisable as a missing binary rather than a broken eval.

No suite changes are needed. Assertions already match Promptfoo's provider-independent command trajectory steps rather than one agent's tool names, and the shared tracing config already declares Claude's Bash alongside Codex's exec_command.

Testing Instructions

  1. Confirm codex is on your PATH (which codex). If it is not, this PR cannot be tested — that is the constraint being documented.
  2. npm run test:agent-evals -- --config specs/testing-skill-routing/test-skill-routing.test.js
  3. npm --workspace @wordpress/agent-skill-evals run view
  4. Confirm the results table has a column per agent, claude and codex, with the same metrics evaluated for both.
  5. npm run test:agent-evals -- --config … --filter-providers codex — confirm a single-agent run works.

Each run makes real model calls against your own quota. Results vary between runs; the agent does not always invoke the skill, so a differing result is expected rather than a broken PR.

Testing Instructions for Keyboard

Not applicable; this PR does not change the user interface.

Screenshots or screencast

Not applicable.

Use of AI Tools

Claude Code (Opus) was used to implement and document this, and Codex (GPT-5) earlier in the series. All output was reviewed and verified by running the harness.

jeryj added 7 commits August 24, 2026 10:51
Adds `test/ai-development/evals`, a standalone workspace that runs coding
agents (Claude and Codex) against isolated temporary checkouts and grades
the result, so we can measure whether the skills in this repository are
discovered and followed.

Promptfoo drives the prompt x provider x test matrix. Lifecycle hooks in
`lib/workspace-extension.mjs` create a clean Git workspace from HEAD before
each row and delete it afterwards; shared provider and assertion defaults
live in `lib/`, and each suite under `specs/` supplies its own cases.

The workspace also regenerates `.claude/skills` from `.agents/skills`. A real
checkout gets that from `npm run agents:setup` via postinstall, but the
generated directory is ignored by Git and the workspace never runs npm — so
without it the repository's skills would sit there as files nothing announces,
and every suite would be measuring a broken environment rather than the
guidance under test.

The first suite covers testing-skill routing. `npm run test:grader`
unit-tests any suite's deterministic grader without spending agent tokens.

Run the evals with `npm run test:agent-evals`.
The Claude provider listed neither the skills to enable nor the tool that
invokes them, so the repository's skills were never offered to the model. Runs
looked like an agent ignoring its guidance when it had never been shown any.

Two settings are needed and they do different jobs. `skills` lists skills to
the model; omitting it is not "skills off" but no SDK configuration at all.
`custom_allowed_tools` replaces the allowed tool list outright rather than
extending it, so `Skill` has to appear there for the listed skills to be
invocable.

With both in place the agent invokes the testing skill and follows it to the
e2e reference, leaving the Jest and PHPUnit references unread.
Two corrections to the testing-skill-routing suite, both found by running it.

The skill assertion matched a shell command reading `SKILL.md`, so it could
only pass if the agent opened the file by hand. Claude invokes the skill
natively instead, which left the assertion failing on a run where the skill was
used correctly. Promptfoo's `skill-used` matches the invocation itself.

The rubric also required asserting serialized content without `expect.poll()`.
Nothing in the testing skill says that, and the contributor documentation
teaches the opposite: `docs/contributors/code/e2e/overusing-snapshots.md`
presents `expect.poll( editor.getBlocks )` as the recommended pattern, and the
suite uses `expect.poll( editor.getEditedPostContent )` 164 times across 14
spec files. Playwright only auto-retries locator assertions, so awaiting the
getter first collapses it to a one-shot comparison — the weaker form of the
two. The clause failed correct work, so it is gone.

Skill invocation is not deterministic: across three completed runs on the fixed
configuration the agent invoked the skill twice. Single runs cannot support
conclusions about this suite.
The Codex provider had never been run. Its binary is not on a normal
non-interactive PATH, so the documented command failed for anyone who tried it,
and an untested provider in the matrix reads as a broken harness rather than an
unfinished one.

Removing it leaves the plain command working with no provider filter. The
agent-rubric grader moves to Claude as well, since it was configured to grade
through Codex.

The shared layout still expects more than one agent — providers live in their
own file and assertions match Promptfoo's provider-independent command
trajectory steps — so a second agent can be added once someone has actually
run it.
The task asked for a test proving a typed `&` stays `&` rather than becoming
`&` in serialized content. That is not how the editor behaves, and should
not be: `escapeAmpersand` in `@wordpress/escape-html` escapes a bare ampersand,
so correct output is `<p>&amp;</p>`. The rubric demanded an assertion that
would fail against correct code, and the agent that noticed and asserted on the
rendered DOM instead was marked wrong for being right.

Centring a paragraph has no such ambiguity. The rubric names the attribute the
editor actually writes — `{"style":{"typography":{"textAlign":"center"}}}` —
and calls out the legacy `{"align":"center"}` form, which is only parsed as
input and is easy to copy from the block fixtures.

Both runs now pass the rubric, leaving skill invocation as the only varying
signal.
The package sat at `test/ai-development/evals`, one level below the `test/*`
workspace glob, so a root `npm install` skipped it and it carried the only
nested `package-lock.json` in the repository. Anyone running the evals had to
discover a second install step first.

Moving it to `test/ai-development` matches the other test workspaces, resolves
it through the root lockfile, and lets the root script address it by name.

The cost is that promptfoo's dependencies now install for every contributor —
about 640 additional entries in the root lockfile. They are devDependencies, so
nothing reaches a published package or plugin build.
The task moved from ampersand encoding to paragraph centre alignment, but the
case kept its old description, which is the label shown in results.
@github-actions

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: jeryj <jeryj@git.wordpress.org>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@github-actions

Copy link
Copy Markdown

Size Change: 0 B

Total Size: 7.9 MB

compressed-size-action

@jeryj
jeryj marked this pull request as draft August 25, 2026 13:41
@jeryj jeryj self-assigned this Aug 25, 2026
@jeryj jeryj added the [Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests. label Aug 25, 2026
jeryj added 2 commits August 25, 2026 09:05
The sandbox options claimed more than they delivered. A probe run showed the
agent could read anything on disk — including this repository's checkout, where
the eval's own assertions and rubrics live. Removing the evaluation directory
from the workspace copy does not help when the original is still readable.

`allowRead` only widens access, and the `allowManaged*Only` locks bind solely
when passed as managedSettings, where they are filtered through a
restrictive-key allowlist; as plain sandbox options they did nothing. So the
source root is now denied explicitly, which a repeat of the probe confirms:
reading a spec file returns "Operation not permitted" rather than its contents.

`failIfUnavailable` is dropped as well — it already defaults to true whenever
the sandbox is enabled.

What remains is verified in force: network egress fails, writes stay inside the
workspace, and the suite still passes.
Three settings were being applied through our own mechanisms rather than the
documented ones.

The per-test timeout lived in an environment variable inside the npm script, so
it was invisible to anyone reading a spec, could not vary per suite, and was
lost when running `promptfoo eval` directly. `evaluateOptions.timeoutMs` does
the same job in the spec, verified by watching a deliberately short value abort
a run with no environment variable set.

`options.bustCache` only suppressed cache reads while still writing entries.
`evaluateOptions.cache` covers both: `initializeAgenticCache` returns early
when caching is disabled, so the per-test option is redundant.

Tracing keeps its `$ref`. Promptfoo resolves JSON schema references through
`@apidevtools/json-schema-ref-parser` and documents them for reusable blocks,
so this is its own sharing mechanism rather than something incidental.
jeryj added 6 commits August 25, 2026 11:41
The extension reproduced what `tools/agents/setup-skills.mjs` does rather than
calling it, so the workspace would silently stop matching a real checkout if
that script ever did more than copy a directory. It exports `setupSkills` and
takes the repository root as an argument, so the workspace can be passed
directly; only its CLI wrapper hardcodes the root.

A missing source directory now fails instead of returning quietly. A workspace
without the generated view still runs, and every suite would report an agent
ignoring guidance it was never offered — the most misleading result this
harness can produce.
The repository writes American English; the task prompt and rubric had crept
into British forms.
Tracing, extensions, providers, defaultTest, evaluateOptions and outputPath are
the same for every suite, but each spec declared its own. A second suite could
quietly run with a different timeout, repeat count or cache setting than the
first, and the difference would not be visible from either file.

They now live in `lib/promptfooconfig.yaml` and are pulled in by reference, so a
spec carries only what is genuinely its own: what it asks for and what it
asserts. Promptfoo resolves JSON schema references, and `file://` paths inside a
referenced value still resolve relative to the spec, so the provider and
defaultTest files load as before.

Results are written to one file per run rather than per suite. Promptfoo's own
store is the history that `npm run view` reads; this is a convenience dump.
Promptfoo accepts `.js` configs alongside `.yaml`, and this repository is
JavaScript nearly everywhere. More usefully, JavaScript composes: a spec spreads
`lib/base.js` and adds only what is its own, where the YAML version repeated a
reference per shared key and could not express partial overrides at all.

The `file://` indirection for providers and default test options goes too —
those are now plain imports.
The YAML carried a `yaml-language-server` schema comment, which editors used to
check the config as it was written. Moving to JavaScript dropped that. Promptfoo
exports its config types, so a JSDoc annotation restores it, matching how the
repository types other configuration — `packages/stylelint-config/index.js` does
the same against stylelint's.

It earns its place immediately: `acceptFormats: [ 'json' ]` widened to
`string[]` where the schema wants `('json' | 'protobuf')[]`, which nothing would
have caught before.
Claude is the only agent in the harness. Running two over the same suite is the
point of the matrix: guidance that works for one agent and not the other is
worth knowing about, and is invisible with a single provider.

No suite changes are needed. Assertions match Promptfoo's provider-independent
command trajectory steps rather than one agent's tool names, and the shared
tracing config already declares Claude's `Bash` alongside Codex's
`exec_command`.

The `codex` binary is not on a normal PATH, so the README says where it comes
from and a failure reads as a missing binary rather than a broken eval.
@jeryj
jeryj force-pushed the add/promptfoo-codex branch from bc763fb to 9c718dc Compare August 25, 2026 18:39
@ciampo
ciampo force-pushed the add/promptfoo branch 8 times, most recently from ac47589 to 059432e Compare August 31, 2026 18:45
Base automatically changed from add/promptfoo to trunk September 4, 2026 21:21

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

[Type] Automated Testing Testing infrastructure changes impacting the execution of end-to-end (E2E) and/or unit tests.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant