Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
9289f4a
docs(spec): add design docs for regression commit attribution
exesilvestre Sep 12, 2026
da08c8e
fix(tests): let AzdStub pass through non-azd subprocess calls
exesilvestre Sep 12, 2026
129f7bc
feat(core): add CommitInfo, ChangedInput, and RegressionInsight models
exesilvestre Sep 12, 2026
6603210
feat(pipeline): resolve the commit an eval run was produced from
exesilvestre Sep 12, 2026
d3399c3
feat(pipeline): attach commit metadata to every evaluation run
exesilvestre Sep 12, 2026
b44777f
feat(pipeline): deterministic diff and regression-insight builder
exesilvestre Sep 12, 2026
e09556d
feat(pipeline): attach a regression insight to --baseline comparisons
exesilvestre Sep 12, 2026
0bd488e
feat(pipeline): render a Regression Insight section in report.md
exesilvestre Sep 12, 2026
d5d78f4
feat(agent): surface regression insight in Doctor's rolling check
exesilvestre Sep 12, 2026
931c02c
feat(cockpit): add an Evaluation Version History section
exesilvestre Sep 12, 2026
b8da9c9
test: end-to-end coverage for regression commit attribution
exesilvestre Sep 12, 2026
1779c8f
docs: document regression commit attribution
exesilvestre Sep 12, 2026
6d309ea
Merge branch 'main' into 012-regression-commit-attribution
exesilvestre Sep 12, 2026
ff3126b
Merge upstream Azure/agentops main (0.15.1) into feature branch
exesilvestre Oct 3, 2026
9b130d9
Merge branch '012-regression-commit-attribution' of https://github.co…
exesilvestre Oct 3, 2026
9b92eff
fix(pipeline): resolve commit metadata from the config's workspace, n…
exesilvestre Oct 8, 2026
ec46e77
fix(pipeline): use the PR head SHA on pull_request events, drop inval…
exesilvestre Oct 8, 2026
71ff7ae
feat(pipeline): list every regressed metric in the insight, not just …
exesilvestre Oct 8, 2026
0346262
fix(cockpit): version history is direction-aware, names regressed met…
exesilvestre Oct 8, 2026
6f88823
fix(doctor): explain unavailable attribution; attribute across versio…
exesilvestre Oct 8, 2026
f0c2a35
fix(doctor,cockpit): fix lineage key for direct-model targets; detect
exesilvestre Oct 8, 2026
58838bb
fix(pipeline): fall back to GITHUB_SHA when PR head isn't resolvable;
exesilvestre Oct 8, 2026
a436164
docs: fix Cockpit eval-history payload contract to match implementation
exesilvestre Oct 8, 2026
2f3313f
fix(cockpit): don't crash the whole page on a non-numeric metric value
exesilvestre Oct 8, 2026
b7f5084
docs: align spec/research/data-model with the shipped version-blind l…
exesilvestre Oct 8, 2026
099557b
fix(doctor): keep rolling-baseline detection on methodology_fingerp…
exesilvestre Oct 8, 2026
6d10045
fix(pipeline): capture commit before execution starts, not after it f…
exesilvestre Oct 8, 2026
8e28f07
fix(doctor): don't fabricate a cause when the attributed pair didn't …
exesilvestre Oct 8, 2026
a282bf4
metrics improvements
exesilvestre Oct 8, 2026
d7035f4
fix(cockpit,doctor): prefer url over name in lineage identity
exesilvestre Oct 8, 2026
ca026eb
fix(pipeline): require same lineage before causal attribution, not ju…
exesilvestre Oct 8, 2026
7fffbfe
fix(pipeline): normalize hosted-agent URLs for lineage; attribute
exesilvestre Oct 8, 2026
d313df4
fix(pipeline): handle boolean threshold criteria; split attribution's
exesilvestre Oct 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,26 @@ This format follows [Keep a Changelog](https://keepachangelog.com/) and adheres

## [Unreleased]

### Added
- **Regression commit attribution.** Every evaluation run now records the
git commit it was produced from (`commit` field in `results.json`) when
it can be determined - reliably for Foundry hosted/prompt agents
evaluated via cloud or azd execution in CI, and on a best-effort basis
for local runs inside a git repository. When a metric regresses between
two comparable runs that both have commit metadata, the evaluation
report gains a "Regression Insight" section explaining, in plain
language, what changed (system prompt, model, dataset, evaluators, or
thresholds) and suggesting a corrective action - surfaced automatically
in the same `report.md` already attached to the PR pipeline, with no new
CI step required. Doctor's rolling-baseline regression check gains the
same explanation in its finding's recommendation. Cockpit gains a new
"Evaluation Version History" section listing every evaluated run with
its commit and what changed relative to the previous run in its
lineage, independent of whether that run regressed. This feature is
purely additive and informational: no new CLI flags, no change to
exit-code/threshold-gating behavior, and existing runs without commit
metadata continue to work exactly as before.

## [0.15.1] - 2026-09-07

### Changed
Expand All @@ -20,6 +40,7 @@ This format follows [Keep a Changelog](https://keepachangelog.com/) and adheres
discover/check workflow bootstraps the identity's profile ID and verifies its
explicit publisher role behind environment approvals, without release uploads.


## [0.15.0] - 2026-09-06

### Added
Expand Down
52 changes: 52 additions & 0 deletions docs/how-it-works.md
Original file line number Diff line number Diff line change
Expand Up @@ -716,6 +716,58 @@ AgentOps writes to both:

If you pass `--output`, AgentOps writes to that directory and still updates `.agentops/results/latest/` with the newest run content.

### Regression commit attribution

Every run's `results.json` also carries a `commit` field (SHA, subject,
author, timestamp) when it can be determined - reliably in CI (via the same
`GITHUB_SHA`/`BUILD_SOURCEVERSION` environment variables already used for
Foundry prompt-agent deploy gating), and on a best-effort basis locally via
`git rev-parse HEAD`. It's `null` when the workspace isn't a git repository
or `git` is unavailable; nothing else about the run changes.

When one or more metrics regress between two runs of the same agent
identity (deliberately independent of version/model/deployment - a change
there is itself one of the things attributed) that both have commit
metadata, a dataset or evaluator-set change between them is attributable
too, same as a prompt/model change: the system never requires an
identical dataset/evaluator set to explain a regression, only that it's
the same agent.

For an explicit `--baseline` comparison, this explanation is a
"Regression Insight" section appended to `report.md`, listing every
metric that regressed (not just the worst one) and explaining what
changed between the two commits (system prompt, model, dataset,
evaluators, or thresholds) with a suggested corrective action. When
either run was published to Foundry (`execution: cloud`, or local
`publish: true`), the section also links out to its Evaluations page.
Doctor's rolling regression check surfaces the same explanation
differently: not in `report.md` at all, but in the triggering `Finding`'s
own `recommendation` text (and `evidence["insight"]`) - Doctor's own
rolling-baseline *detection* additionally stays on the stricter,
version-inclusive methodology fingerprint unchanged by any of this (see
`agent.checks.regression`), so a version bump alone doesn't affect
*whether* Doctor flags a regression, only how the already-flagged one
gets explained. This is deterministic field-diffing, not an LLM call, and
it's purely informational: it never affects the exit-code/threshold-gating
contract. Runs without commit metadata, or where nothing tracked changed,
behave exactly as they did before this existed.

Every run produced by `agentops eval run` - including `execution: cloud`
and `execution: azd` - attempts commit capture. The one case with no commit
to attribute is Doctor's rolling regression check when it falls back to
Foundry cloud evaluation runs because local history is too short (see
`results_history`'s cloud fallback): those runs have no local
`results.json` to reload and therefore no recorded commit. When that
happens, the regression finding's recommendation says so explicitly (e.g.
`attribution unavailable: run has no commit (fetched from cloud)`) instead
of silently omitting the insight.

Cockpit's dashboard also gains an "Evaluation Version History" section
listing every locally recorded run, newest first, with its commit, which
metrics (if any) regressed relative to the previous run in its lineage by
name, and what changed - shown regardless of whether that run regressed, so
you can browse the causal trail over time.

## Testing

Tests live in `tests/` and are organized as:
Expand Down
35 changes: 35 additions & 0 deletions specs/012-regression-commit-attribution/checklists/requirements.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Specification Quality Checklist: Regression Commit Attribution

**Purpose**: Validate specification completeness and quality before proceeding to planning
**Created**: 2026-09-11
**Feature**: [spec.md](../spec.md)

## Content Quality

- [x] No implementation details (languages, frameworks, APIs)
- [x] Focused on user value and business needs
- [x] Written for non-technical stakeholders
- [x] All mandatory sections completed

## Requirement Completeness

- [x] No [NEEDS CLARIFICATION] markers remain
- [x] Requirements are testable and unambiguous
- [x] Success criteria are measurable
- [x] Success criteria are technology-agnostic (no implementation details)
- [x] All acceptance scenarios are defined
- [x] Edge cases are identified
- [x] Scope is clearly bounded
- [x] Dependencies and assumptions identified

## Feature Readiness

- [x] All functional requirements have clear acceptance criteria
- [x] User scenarios cover primary flows
- [x] Feature meets measurable outcomes defined in Success Criteria
- [x] No implementation details leak into specification

## Notes

- Reasonable defaults (reuse of existing methodology-fingerprint comparability, existing regression-detection thresholds, CI-provided commit SHA availability) are documented in the Assumptions section rather than raised as clarification questions, per product direction already given (prioritize the hosted-agent/CI scenario) in the source conversation.
- All items pass; no follow-up needed before `/speckit-plan`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# Contract: `report.md` section and Cockpit history payload

## `report.md`: new "Regression Insight" section

Rendered by `pipeline/reporter.py` immediately after the existing "Comparison
vs Baseline" section (`_render_comparison`), and only when
`result.comparison.insight` is present:

```markdown
## Comparison vs Baseline
...existing table...

## Regression Insight

**Regressed metrics:**
- `similarity`: 4.000 → 3.000
- `accuracy`: 0.910 → 0.790

Run v3 → v4: similarity dropped from 4.00 to 3.00 and accuracy dropped from
0.91 to 0.79. Likely cause: the system prompt changed and the model changed
from `gpt-4o` to `gpt-4o-mini`.

**Suggested action:** review the prompt change and the model swap; consider
reverting one at a time to isolate the cause.

- [Baseline run in Foundry](https://ai.azure.com/...)
- [Regressed run in Foundry](https://ai.azure.com/...)
```

Every metric that regressed is listed - not just the single worst one - in
`RegressionInsight.regressed_metrics`, ordered worst-first
(direction-aware, so a lower-is-better metric like `avg_latency_seconds`
getting worse still sorts as a regression). The "Regressed metrics" bullet
list only renders when there's more than one; a single-metric regression
keeps the original one-line prose shape. The two Foundry links
(`from_report_url`/`to_report_url`) render only when that side was
published (`execution: cloud`, or a completed local `publish: true`) -
omitted silently otherwise, never fabricated. For the live `--baseline`
comparison specifically, the *current* run's own link is initially `None`
when the run will go on to `publish: true`, because that publish step
happens after this comparison is built and persisted - but
`orchestrator._publish_to_foundry_safely` patches the value in and
re-persists `results.json`/`report.md` once that publish step actually
completes, so the already-written files don't permanently miss their own
link. Doctor's rolling check, which compares two already-completed
historical runs, never has this gap in the first place.

This section is purely additive to the report: it never replaces the
existing Metrics/Thresholds/Comparison/Rows sections, and its absence (no
regression, or missing commit metadata on either side) leaves `report.md`
byte-for-byte identical to today's output. Because this is the same
`report.md` already uploaded as the `agentops-pr-results` CI artifact and
posted as the PR comment by the generated `agentops-pr.yml` workflow, no
workflow template changes are required to deliver it (FR-010).

Doctor's rolling-baseline `Finding` (`agent/checks/regression.py`) gets the
same explanation text via `Finding.evidence["insight"]`, and Doctor's own
Markdown rendering already surfaces `Finding.summary`/`recommendation` — this
plan extends that finding's `recommendation` text with the same
deterministic explanation when an insight is available, in place of today's
generic "inspect prompt/model/dataset changes" instruction. When no insight
could be built at all (e.g. Doctor fell back to a Foundry cloud-fetched run
with no local `results.json`, so no commit), the recommendation instead
names why attribution is unavailable
(`Finding.evidence["attribution_unavailable"]`) rather than silently
omitting it.

## Cockpit: version history

No new HTTP route. The existing partial-load endpoint
(`GET /?_partial=1`) response gains a new `eval_history` section
(`_build_eval_history_section`), alongside the existing eval "cards"
section. This is the actual payload shape (`cockpit.py`'s
`_build_eval_history_section`/`_attach_version_history`), **newest run
first**:

```jsonc
{
// ...existing cockpit payload sections unchanged...
"eval_history": {
"has_runs": true,
"entries": [
{
"run_id": "20260910-140300",
"timestamp": "2026-09-10T14:03:00Z",
"target": "greeter:4",
"metrics": { "accuracy": 0.79 },
"commit_short_sha": "b7e91aa",
"commit_subject": "Swap eval agent to gpt-4o-mini",
"changed_inputs": [
{ "field": "model", "description": "model changed from gpt-4o to gpt-4o-mini", "from_value": "gpt-4o", "to_value": "gpt-4o-mini" }
],
"regressed": true,
"regressed_metrics": ["accuracy"],
"report_link": "/api/runs/20260910-140300/report",
"cloud_report_url": null,
"previous_cloud_report_url": "https://ai.azure.com/foundry/baseline"
},
{
"run_id": "20260901-101500",
"timestamp": "2026-09-01T10:15:00Z",
"target": "greeter:3",
"metrics": { "accuracy": 0.91 },
"commit_short_sha": "a1b2c3d",
"commit_subject": "Initial smoke eval",
"changed_inputs": [],
"regressed": false,
"regressed_metrics": [],
"report_link": "https://ai.azure.com/foundry/baseline",
"cloud_report_url": "https://ai.azure.com/foundry/baseline",
"previous_cloud_report_url": null
}
]
}
}
```

Entries are grouped/compared using `version_lineage_key` - a coarser,
version/deployment-blind key (same agent identity, dataset, evaluators)
than Doctor's `methodology_fingerprint`, deliberately: this view exists to
show what changed *across* version bumps, where Doctor's rolling check
deliberately excludes a version bump from its own baseline. `commit` is
flattened to `commit_short_sha`/`commit_subject` (not a nested object); a
run with no resolvable commit still appears, with both `null` and an
empty `changed_inputs` list rather than being omitted (per User Story 2's
acceptance scenario 3). `regressed_metrics` names every metric that
regressed relative to the previous entry in its lineage, not just a
boolean - a lower-is-better metric (e.g. `avg_latency_seconds`) getting
*better* is never counted as a regression. `cloud_report_url` (this run's
own Foundry Evaluations link, `null` if never published) and
`previous_cloud_report_url` (the previous lineage-comparable run's, same
rule) let a regressed row link to both sides of the comparison; `report_link`
is unrelated pre-existing per-row routing (Foundry when published, else
this run's own local report page) and always non-null. This view is
read-only and introduces no writes to any monitored resource, consistent
with the constitution's Cockpit read-only requirement.
76 changes: 76 additions & 0 deletions specs/012-regression-commit-attribution/contracts/results-json.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
# Contract: `results.json` additions

`results.json` is a documented public contract (constitution, Principle I).
This feature only adds optional, additive fields — no existing field
changes meaning or type, and no consumer that ignores unknown-to-it new
top-level keys breaks.

## New top-level field: `commit`

```jsonc
{
// ...existing RunResult fields unchanged...
"commit": {
"sha": "a1b2c3d4e5f6...",
"short_sha": "a1b2c3d",
"subject": "Swap eval agent to gpt-4o-mini",
"author": "Jane Doe",
"authored_at": "2026-09-10T14:03:00Z",
"source": "ci"
}
// or "commit": null when it could not be determined
}
```

Absent/`null` on any run produced before this feature shipped, or when commit
metadata could not be resolved (not a git repo, git unavailable). Consumers
MUST treat a missing/`null` `commit` the same as before this feature existed.

## New field on `comparison`: `insight`

Only present when `comparison` (the existing `--baseline` block) is present
**and** at least one metric regressed **and** both the current and baseline
runs have non-null `commit`. `regressed_metrics` lists every metric that
regressed, not just the worst one, ordered worst-first:

```jsonc
{
"comparison": {
// ...existing ComparisonInfo fields unchanged...
"insight": {
"from_run_id": "20260901-101500",
"to_run_id": "20260910-140300",
"from_commit": { "sha": "...", "short_sha": "...", "subject": "...", "author": "...", "authored_at": "...", "source": "ci" },
"to_commit": { "sha": "...", "short_sha": "...", "subject": "...", "author": "...", "authored_at": "...", "source": "ci" },
"regressed_metrics": [
{ "metric": "accuracy", "from_value": 0.91, "to_value": 0.79 }
],
"changed_inputs": [
{ "field": "system_prompt", "description": "the system prompt changed", "from_value": "greeter:v3", "to_value": "greeter:v4" },
{ "field": "model", "description": "model changed from gpt-4o to gpt-4o-mini", "from_value": "gpt-4o", "to_value": "gpt-4o-mini" }
],
"explanation": "Run v3 → v4: accuracy dropped from 0.91 to 0.79. Likely cause: the system prompt changed and the model changed from gpt-4o to gpt-4o-mini.",
"suggested_action": "Review the prompt change and the model swap; consider reverting one at a time to isolate the cause.",
"commits_available_locally": false,
"from_report_url": null,
"to_report_url": null
}
}
}
```

Serialized as `comparison.insight: null` (the key is present, since
`model_dump(mode="json")` is called without `exclude_none`) whenever no
metric regressed or either compared run lacks commit metadata. When
`comparison` itself is absent (no `--baseline` was used), there is no
nested `insight` field at all. Either way, every existing
`comparison`-shaped consumer that ignores unknown-to-it or null fields is
unaffected.

## No changes to exit codes or CLI flags

This feature introduces no new `agentops eval run` flags and does not change
the exit-code contract (`0`/`1`/`2` meanings are unchanged, per FR-011). The
committed-baseline auto-detection already used by generated PR workflows
(`.agentops/baseline/results.json` → `--baseline`) is unchanged; this feature
only adds richer content to what that path already produces.
Loading