You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Right now, when an eval regresses, there's no way to know why without manually digging through what changed between that run and the previous one. This PR adds that: every run now records the commit it came from, and if a metric regresses between two comparable runs (via --baseline or Doctor's check), the report explains in one sentence what changed (prompt, model, dataset, evaluators) and suggests a fix. All deterministic, no LLM involved.
Prioritizes the hosted-agent-in-CI scenario (cloud/azd), since that's where partial commit info was already available.
Also adds a version-history view to Cockpit, showing what changed between each evaluated version, regression or not.
Full design in specs/012-regression-commit-attribution/ if you want the details.
Relates to #494. That issue proposed something broader (a separate Doctor check, a real git diff with line counts, a dedicated Cockpit endpoint). Here the scope is smaller: the insight only fires when there's already been a regression, the diff is over fields already tracked in results.json (no actual file git diff), and Cockpit reuses the existing payload instead of a new endpoint. Still lines up on the core idea: group by dataset+evaluators without the target version, so you can compare across a prompt/model change.
Test plan
131 tests across unit and integration (commit capture local/CI, the diffing, wiring into --baseline and Doctor, the rendering, Cockpit, and an end-to-end case with a real server).
Manually tested with the real CLI against a real git repo.
Also validated against a real Foundry hosted agent in Azure: a version bump of the agent's prompt was detected as a regression with the correct commit attribution in report.md and Cockpit.
Hi Exequiel SIlvestre (@exesilvestre), thanks for this PR. Review notes below, each with a suggested fix. Items 5 and 6 are suspected from reading the diff and not yet run, so please confirm them.
1. Git runs in the wrong folder
resolve_commit_info() and commit_exists_locally() accept workspace, but callers never pass it, so git uses the process cwd.
Example: from C:\work\other-project, agentops eval run --config C:\work\my-agent\agentops.yaml records other-project's HEAD. Fix: pass workspace=config_path.parent in _finalize_commit_and_comparison and _persist. Add an optional workspace param to build_comparison and build_regression_insight. Add a test that runs from a different cwd.
2. GITHUB_SHA is the wrong commit on pull_request events
On pull_request events it is the temporary merge commit, not the PR head. Fix: when GITHUB_EVENT_NAME == "pull_request", read pull_request.head.sha from GITHUB_EVENT_PATH; then fall back to GITHUB_SHA, then git rev-parse HEAD. Add tests.
3. Invalid env var name Build.SourceVersion
In _CI_SHA_ENV_VARS, the real env var is BUILD_SOURCEVERSION. Fix: remove the invalid name and add a test for BUILD_SOURCEVERSION.
4. used_git_diff does not run a diff
It only checks that both commits exist locally. Fix: rename to commits_available_locally, or run git diff --name-only from..to. Document that prompt edits without a version bump go undetected.
5. (please confirm) Cockpit regressed seems to ignore metric direction
Lower-is-better metrics (e.g. latency) look to be treated as higher-is-better. Fix: reuse the direction helper from pipeline/comparison.py. Add a latency test.
6. (please confirm) _relative_drop sign and zero baseline
It seems to have the wrong sign for lower-is-better metrics and returns 0.0 when baseline <= 0. Fix: make it direction-aware; use the absolute change when baseline <= 0. Add a test.
7. Conflicting docstrings
_version_lineage_key and methodology_fingerprint are described inconsistently. Fix: align the docstrings, or import methodology_fingerprint instead of duplicating.
8. Repeated git calls and validation cost
_persist calls git a second time after a failed first attempt (5s timeout each). Cockpit validates up to 24 RunResults per render. Fix: use a flag to skip the second git call; cache validation by (path, mtime).
9. Local vs cloud runs in Doctor
Runs from agentops eval run always record commit, including execution: cloud and azd. But Doctor also merges runs fetched from Foundry (_merge_runs fallback). Those have no commit and no loadable RunResult, so the insight is silently skipped when either side is such a run. Fix: show "attribution unavailable: run has no commit (fetched from cloud)" when skipped. Document it in the docs and doctor explain. Add a test with a local baseline and a cloud-fetched current run.
10. Only one regressed metric is explained
build_comparison picks the single worst regressed metric, so with several regressions (e.g. similarity, coherence, latency) the others get no mention in the insight, report or Cockpit. Doctor already reports one finding per metric, so the two surfaces differ. Fix: let the insight list all regressed metrics (e.g. regressed_metrics: [{metric, from, to}]), ordered direction-aware. The cause (changed_inputs) is shared. Render all of them in the report and Cockpit. Add a test with 3 regressed metrics.
11. Link to the cloud evaluation
Runs with execution: cloud (or publish: true) already write cloud_evaluation.json with report_url. Fix: include the report_url of the from and to runs in the insight, report and Cockpit when present; omit silently when absent. Informational only: it must not affect thresholds or exit codes. Add a test for both cases.
This call runs two git cat-file subprocesses inside the per-metric loop. With the six default watched metrics, one Doctor analysis can repeat the identical commit-availability lookup up to twelve times, and each lookup has a five-second timeout. Compute commit availability once for the run pair and reuse it across all metric insights.
Leave changed_inputs empty when commit metadata is missing
src/agentops/agent/cockpit.py:998
This computes and displays changed inputs even when either run has no commit metadata. The feature contract explicitly requires an older/unknown-commit history entry to have an empty changed_inputs list rather than a causal description. Only build the change list when both parsed runs have commits; metric trend calculation can remain independent.
Honor configured lower-is-better metric direction
src/agentops/pipeline/regression_insight.py:50
Lower-is-better direction is configurable for any metric: Threshold.from_expression accepts < and <=, not only avg_latency_seconds. A custom metric such as error_rate: "<=0.1" will therefore treat an increase as an improvement, suppressing both the comparison insight and Cockpit regression badge. Derive direction from the run's recorded threshold criteria, using higher-is-better only when no directional criterion exists.
Exclude metrics that did not regress from the prior run
The supplied metric is appended without confirming that it regressed between these two runs. Doctor detects a drop against the rolling mean, but from_run is only the immediately preceding run; for example, a current 0.75 versus prior 0.70 and rolling baseline 0.90 produces an insight calling an improvement a regression and attributing a false likely cause. Filter out metrics that are not actually worse than from_run.
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
The cached projection also depends on the sibling cloud_evaluation.json, but only results.json contributes to this cache key. During a local publish: true run, Cockpit can cache the result before publishing finishes; creation of the sidecar then leaves the cached cloud_report_url and Foundry history links stale for the lifetime of the server. Include sidecar existence/mtime (preferably an (mtime_ns, size) signature for both files) in the cache key.
Clarify insights remain when no tracked inputs changed
docs/how-it-works.md:740
This says runs with no tracked input changes behave as before, but build_regression_insight deliberately returns an insight with the metric/commit explanation and the reporter renders the new section even when changed_inputs == [] (also asserted by test_insight_explanation_has_no_cause_when_nothing_tracked_changed). Clarify that only the likely-cause text is omitted in this case.
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
Require commit metadata before attributing changed inputs
src/agentops/agent/cockpit.py:1001
The version-history contract says a run without recorded commit metadata must have an empty changed_inputs list, but this comparison runs for fully valid legacy results whose commit defaults to None. Consequently a no-commit run following an older run displays an attributed change description, contrary to the documented graceful-fallback behavior. Gate change attribution on both parsed runs having commits while still allowing the metric regression calculation below.
Invalidate cache when cloud evaluation sidecar changes
src/agentops/agent/cockpit.py:1062
The cache signature only tracks results.json, but _project_run_uncached() also reads the sibling cloud_evaluation.json. For a running Cockpit, a local publish normally writes results.json first and the sidecar later; when there is no regression insight, publishing does not rewrite results.json, so a projection cached during that window will keep cloud_report_url=None for the lifetime of the server. Include the sidecar's existence/mtime in the cache signature (and invalidate when it changes).
Exclude threshold edits from causal changes and recommendations
Every tracked difference is presented as a likely cause, including thresholds. Threshold criteria are applied after scores are produced and cannot cause a raw metric value to drop, so a run with a threshold edit plus a stochastic score drop gets a false causal statement and a suggestion to revert an irrelevant change. Keep threshold edits as contextual changes, but exclude them from the cause and corrective-action lists.
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
A failed pre-execution lookup is retried here because commit=None is indistinguishable from “the caller did not pre-capture.” Every real local/cloud/azd path passes the failed None, so runs without resolvable commit metadata execute the git probes twice despite the new retry-avoidance intent (the tests only call this helper once and miss that path). Use a sentinel or an explicit commit_resolution_attempted flag so an attempted None is preserved without another lookup.
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
Include cloud report sidecar changes in cache invalidation
src/agentops/agent/cockpit.py:1075
The cache is invalidated only by results.json mtime, although _project_run_uncached also reads the sibling cloud_evaluation.json. If Cockpit scans after results persistence but before cloud publishing creates that sidecar, it caches cloud_report_url=None indefinitely; ordinary published runs do not rewrite results.json. Include the sidecar's existence/mtime in the cache signature.
Exclude result rows from cached version-history projections
src/agentops/agent/cockpit.py:1124
The app-scoped cache retains this full validated RunResult, including every row and response, for up to 24 runs even though version-history comparison only needs target/config/metrics/thresholds. Large histories can therefore keep many complete result payloads resident for the Cockpit process. Drop rows before placing the model in the cached projection.
changed_inputs is populated whenever both results validate, even if either run has no commit metadata. This contradicts the version-history contract for older/no-commit runs, which requires an empty change list rather than attributing changes without a commit pair. Gate only the field diff on both commits; metric trend calculation can remain independent.
Cache key omits cloud_evaluation.json changes
src/agentops/agent/cockpit.py:1065
The cache fingerprint only tracks results.json, but _project_run_uncached also reads the sibling cloud_evaluation.json. A local publish: true run writes that sidecar after the initial persist and does not rewrite results.json unless an insight exists, so Cockpit can cache the pre-publish projection and keep showing the local link for the lifetime of the server. Include the sidecar's presence/mtime in the cache key.
The acknowledged project-blind identity makes the comparability guard unsafe: an explicit baseline from another Foundry project with the same prompt-agent name is treated as the same agent, so the report can attribute an unrelated metric difference to a version/model change. Direct-model runs have the same issue because every model_direct target collapses to one identity. Persist a project-qualified identity (for example the resolved project endpoint/resource ID) and include it here before producing causal attribution.
RegressionInsight should not render without changed inputs
This still constructs and renders a RegressionInsight when changed_inputs is empty. That conflicts with SC-005 and the user-facing documentation, which say a run whose cause cannot be determined keeps the prior report structure; the new section merely repeats metric values and provides neither a cause nor a corrective action. Either return None when no tracked change exists, or explicitly revise the specification and documentation to define a metric-only insight.
The reason will be displayed to describe this comment to others. Learn more.
Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.
This branch has not been deployed
No deployments
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Right now, when an eval regresses, there's no way to know why without manually digging through what changed between that run and the previous one. This PR adds that: every run now records the commit it came from, and if a metric regresses between two comparable runs (via
--baselineor Doctor's check), the report explains in one sentence what changed (prompt, model, dataset, evaluators) and suggests a fix. All deterministic, no LLM involved.Prioritizes the hosted-agent-in-CI scenario (cloud/azd), since that's where partial commit info was already available.
Also adds a version-history view to Cockpit, showing what changed between each evaluated version, regression or not.
Full design in
specs/012-regression-commit-attribution/if you want the details.Relates to #494. That issue proposed something broader (a separate Doctor check, a real
git diffwith line counts, a dedicated Cockpit endpoint). Here the scope is smaller: the insight only fires when there's already been a regression, the diff is over fields already tracked inresults.json(no actual filegit diff), and Cockpit reuses the existing payload instead of a new endpoint. Still lines up on the core idea: group by dataset+evaluators without the target version, so you can compare across a prompt/model change.Test plan
--baselineand Doctor, the rendering, Cockpit, and an end-to-end case with a real server).report.mdand Cockpit.