Repository navigation
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Rescheduling can bypass freshness checks, mutations can proceed without durable audit records, and credential fallback can target unintended databases.
Review effort: Balanced
Findings: 3
Open (9)
Preserve non-NotFound errors as AttestationUnknown · New Prevent legacy database fallback for verifier connections · New Run archive and attestation safety checks on every execution · New Source verifier variable from recovery collection metrics · New Reject negative discovery replay cursors · New Log replay intent before launching the goroutine · New Log recovery intent before creating the operation · New Record action intent before changing operation state · New Persist action intent before executing target mutations · New
What changed in this PR
Adds a server-rendered verifier admin console for cross-node search, recovery, rescheduling, indexer backfills, and audit logging, alongside recovery observability.
Changes:
- Adds the admin CLI, web server, views, migrations, and operational workflows.
- Adds attestation gates, archive/recovery metrics, dashboards, and alerts.
- Expands tests, documentation, and runbooks.
| File | Description |
|---|---|
.github/actions/build-cl/action.yaml |
Updates Chainlink revision. |
.gitignore |
Ignores demo runtime files. |
build/devenv/dashboards/verifier_archive_inventory.json |
Adds archive dashboard. |
build/devenv/dashboards/verifier_recovery.json |
Adds recovery dashboard. |
build/devenv/go.sum |
Updates devenv dependencies. |
build/devenv/tests/e2e/smoke_recovery_cli_test.go |
Adds inventory E2E coverage. |
changelog/2026-09-10_archive_inventory_and_cli.md |
Documents archive features. |
changelog/2026-09-11_source_recovery.md |
Updates recovery changelog. |
changelog/2026-09-23_admin_console.md |
Documents admin console. |
cli/admin/commands.go |
Adds admin CLI commands. |
cli/recovery/README.md |
Links the console workflow. |
cmd/verifier/run_ccv_cli.go |
Registers admin commands. |
docs/config/README.md |
Registers console configuration. |
docs/config/admin-console/config.documented.toml |
Documents console config. |
docs/monitoring/verifier-archive-inventory-alerts.yaml |
Adds archive alerts. |
docs/monitoring/verifier-archive-inventory.md |
Documents archive metrics. |
docs/monitoring/verifier-recovery-alerts.yaml |
Adds recovery alerts. |
docs/monitoring/verifier-recovery.md |
Documents recovery metrics. |
docs/runbooks/remediating-stuck-or-dropped-messages.md |
Adds console procedures. |
docs/verifier/admin-console.md |
Documents console operation. |
go.mod |
Adds templ dependency. |
go.sum |
Records templ checksums. |
integration/storageaccess/aggregator_results_client.go |
Adds aggregator results client. |
integration/storageaccess/aggregator_results_client_test.go |
Tests protobuf translation. |
verifier/pkg/admin/actionlog.go |
Implements action persistence. |
verifier/pkg/admin/actionlog_test.go |
Tests action-log migrations. |
verifier/pkg/admin/attestation.go |
Implements freshness checks. |
verifier/pkg/admin/attestation_test.go |
Tests attestation states. |
verifier/pkg/admin/backfill.go |
Implements indexer backfills. |
verifier/pkg/admin/backfill_test.go |
Tests backfill workflows. |
verifier/pkg/admin/config.go |
Defines console configuration. |
verifier/pkg/admin/config_test.go |
Tests config validation. |
verifier/pkg/admin/db.go |
Opens and migrates databases. |
verifier/pkg/admin/detail.go |
Implements message details. |
verifier/pkg/admin/detail_test.go |
Tests detail rendering. |
verifier/pkg/admin/handlers.go |
Provides shared HTTP handlers. |
verifier/pkg/admin/migrations/embed.go |
Embeds admin migrations. |
verifier/pkg/admin/migrations/postgres/00001_admin_actions.sql |
Creates action-log schema. |
verifier/pkg/admin/node.go |
Encapsulates node connections. |
verifier/pkg/admin/recoveryops.go |
Implements recovery operations. |
verifier/pkg/admin/recoveryops_test.go |
Tests recovery workflows. |
verifier/pkg/admin/reschedule.go |
Implements rescheduling. |
verifier/pkg/admin/reschedule_test.go |
Tests rescheduling behavior. |
verifier/pkg/admin/search.go |
Implements cross-node search. |
verifier/pkg/admin/server.go |
Configures the HTTP server. |
verifier/pkg/admin/server_test.go |
Tests middleware and routes. |
verifier/pkg/admin/views/actions.templ |
Defines action-log view. |
verifier/pkg/admin/views/actions_templ.go |
Generates action-log rendering. |
verifier/pkg/admin/views/backfill.templ |
Defines backfill views. |
verifier/pkg/admin/views/backfill_templ.go |
Generates backfill rendering. |
verifier/pkg/admin/views/detail.templ |
Defines detail view. |
verifier/pkg/admin/views/detail_templ.go |
Generates detail rendering. |
verifier/pkg/admin/views/layout.templ |
Defines shared layout. |
verifier/pkg/admin/views/layout_templ.go |
Generates layout rendering. |
verifier/pkg/admin/views/nodes.templ |
Defines node overview. |
verifier/pkg/admin/views/nodes_templ.go |
Generates node rendering. |
verifier/pkg/admin/views/recovery.templ |
Defines recovery views. |
verifier/pkg/admin/views/recovery_templ.go |
Generates recovery rendering. |
verifier/pkg/admin/views/reschedule.templ |
Defines reschedule views. |
verifier/pkg/admin/views/reschedule_templ.go |
Generates reschedule rendering. |
verifier/pkg/admin/views/search.templ |
Defines search views. |
verifier/pkg/admin/views/search_templ.go |
Generates search rendering. |
verifier/pkg/admin/views/static.go |
Embeds browser assets. |
verifier/pkg/admin/views/static/admin.css |
Styles the console. |
verifier/pkg/admin/views/static/favicon.svg |
Adds console icon. |
verifier/pkg/admin/views/static/htmx.min.js |
Vendors htmx. |
verifier/pkg/jobqueue/archive.go |
Adds archive inventory metrics. |
verifier/pkg/jobqueue/archive_test.go |
Tests inventory lifecycle. |
verifier/pkg/jobqueue/observability_decorator.go |
Schedules archive collection. |
verifier/pkg/jobqueue/postgres_queue.go |
Initializes archive metrics. |
verifier/pkg/jobqueue/testdata/explain_cleanup.txt |
Refreshes cleanup plan. |
verifier/pkg/jobqueue/testdata/explain_complete.txt |
Refreshes completion plan. |
verifier/pkg/jobqueue/testdata/explain_consume_pending.txt |
Refreshes pending plan. |
verifier/pkg/jobqueue/testdata/explain_consume_stale.txt |
Refreshes stale plan. |
verifier/pkg/jobqueue/testdata/explain_fail.txt |
Refreshes failure plan. |
verifier/pkg/jobqueue/testdata/explain_publish_conflict.txt |
Refreshes conflict plan. |
verifier/pkg/jobqueue/testdata/explain_publish_no_conflict.txt |
Refreshes insertion plan. |
verifier/pkg/jobqueue/testdata/explain_retry.txt |
Refreshes retry plan. |
verifier/pkg/jobqueue/testdata/explain_size.txt |
Refreshes size plan. |
verifier/pkg/recovery/metrics.go |
Adds recovery metrics. |
verifier/pkg/sourcereader/recovery.go |
Collects recovery metrics. |
verifier/pkg/sourcereader/recovery_audit.go |
Counts audit failures. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if entries[i].ErrorCode != int32(codes.OK) { | ||
| return AttestationResult{AttestationNotFound, "aggregator: " + entries[i].ErrorMsg} | ||
| } |
There was a problem hiding this comment.
Already fixed on the current head: only per-entry NotFound maps to AttestationNotFound; any other non-OK code is AttestationUnknown, which keeps execution disabled (see aggregatorEntryResult).
| url := secrets.DatabaseURL() | ||
| if url == "" { | ||
| return nil, fmt.Errorf("node secrets file %q has no [db].url", secretsPath) |
There was a problem hiding this comment.
Moot after the rework: the console no longer reads DB URLs from secrets files at all — it shares the verifier's in-process application DB handle, so there is no CL_DATABASE_URL fallback path (and DatabaseURLFileOnly is removed). #1500
| if retryMode { | ||
| state, detail, _ := recheckArchiveRow(ctx, store, t) | ||
| if state == recheckExecutable { | ||
| att := checkNodeAttestations(ctx, n.Config(), [][]byte{t.MessageID})[0] | ||
| switch att.State { | ||
| case AttestationAttested: | ||
| state, detail = recheckSkip, "already attested — nothing to do ("+att.Detail+")" | ||
| case AttestationUnknown: | ||
| state, detail = recheckUnknown, "attestation state unknown: "+att.Detail | ||
| } | ||
| } | ||
| switch state { | ||
| case recheckSkip: | ||
| out.outcome, out.detail = "skipped", detail | ||
| return out | ||
| case recheckUnknown: | ||
| out.outcome, out.detail = "skipped", "not executed — "+detail | ||
| return out | ||
| } | ||
| } |
There was a problem hiding this comment.
Already fixed on the current head: executeTarget re-runs the archive + attestation gate on every execution (direct POST included); TestRescheduleExecuteDirectPostStillGated covers it.
| "definition": "label_values(verifier_archive_collection_success, verifier_id)", | ||
| "query": "label_values(verifier_archive_collection_success, verifier_id)", |
There was a problem hiding this comment.
The recovery dashboard was removed from this PR's scope, so there is nothing to fix here anymore.
| since, err := strconv.ParseInt(sinceStr, 10, 64) | ||
| if err != nil { | ||
| return replay.Request{}, fmt.Errorf("aggregator sequence number must be an unsigned decimal integer: %w", err) | ||
| } | ||
| return replay.Request{Type: replay.TypeDiscovery, Since: since, Force: force}, nil |
There was a problem hiding this comment.
The backfill feature was removed from this PR's scope (rescoped to verifier recovery only), so this path no longer exists.
| if logErr := h.recordAction(c, Action{ | ||
| Action: "backfill-submit", NodeName: res.NodeName, Target: target, | ||
| OperationID: res.JobID, Outcome: "success", Detail: detail, | ||
| }); logErr != nil { | ||
| res.Error = "replay job was started but the action log write failed: " + logErr.Error() |
There was a problem hiding this comment.
Same — the backfill goroutine path was removed from this PR's scope.
| if err := h.recordAction(c, Action{ | ||
| Action: "recovery-submit", NodeName: name, Target: target, | ||
| OperationID: op.ID, Outcome: "success", Detail: detail, | ||
| }); err != nil { | ||
| res.Error = "operation " + op.ID + " was created but the action log write failed: " + err.Error() |
There was a problem hiding this comment.
Already fixed on the current head: the intent row (outcome=started) is written before the operation is created; an unloggable submission never reaches the store.
| if logErr := h.recordAction(c, Action{ | ||
| Action: "recovery-" + action, NodeName: nodeName, Target: target, | ||
| OperationID: id, Outcome: outcome, Detail: detail, | ||
| }); logErr != nil { | ||
| detail += " (action log write failed: " + logErr.Error() + ")" |
There was a problem hiding this comment.
Already fixed on the current head: intent is logged before ChangeState, outcome after.
| if err := h.recordAction(c, Action{ | ||
| Action: "reschedule", NodeName: o.target.NodeName, Target: logTarget, | ||
| Outcome: o.outcome, Detail: o.detail, | ||
| }); err != nil { | ||
| auditErrs = append(auditErrs, fmt.Sprintf("%s: %v", logTarget, err)) |
There was a problem hiding this comment.
Already fixed on the current head: executeTarget writes the intent row before touching the queue, and an action-log failure aborts the mutation.
makramkd
left a comment
There was a problem hiding this comment.
I think there's a few simplifications we can do, listed some out in the comments, but overall:
- I think we can assume, like the chainlink operator UI, that this admin UI is only ever going to be connected to a single verifier node. As such I think we can probably simplify config (I don't think we need the
Nodesfield for example) and I think we can look into embed the admin config in the full verifier config (might be a bit difficult if it references values only the deployer would know, e.g. paths to secrets files - if this is the case a separate admin config file is fine IMO). - I think we can get rid of
ccv admin serveas a CLI - if nops want the admin UI they should add the required secret (uname/pw) and admin config file and just restart the pod, in which case it'll launch it automatically. - Migrations can probably be folded into existing verifier migrations.
Maybe in the future there's some re-use across admin UIs for the different apps, but as a starting point I think w/ these simplifications it'd be a good ship. Open to your thoughts on them.
| go func() { | ||
| defer close(exited) | ||
| for { | ||
| child := exec.Command(exe, "ccv", "admin", "serve", "--config", path) //nolint:gosec // G204: re-exec of this binary with fixed argv; the config path is the operator's own deployment. |
There was a problem hiding this comment.
Does this have to be a literally separate process, or can we call admin.NewServer directly and serve on a pre-configured/constant port (e.g. 9099 or whatever)?
I'm not sure about the separate ccv admin serve use-case, are you thinking that that command can be used if the NOP didn't provide the required config/secrets from the get-go?
I think that might simplify a lot of the process-related code here.
There was a problem hiding this comment.
Done — the sibling process is gone. The factories call admin.NewServer directly and serve in-process on the configured listen_address (default 127.0.0.1:8105), so all the process-management code is deleted. ccv admin serve is removed; enabling the console = add the config file (+ [admin_ui] in the verifier secrets if wanted) and restart. See #1501.
| // A present console config serves the admin UI as a sibling process in this | ||
| // container; the verifier's own lifecycle is unaffected. | ||
| if stopConsole := cmd.StartAdminConsoleSibling(); stopConsole != nil { | ||
| defer stopConsole() | ||
| } |
There was a problem hiding this comment.
I wonder if this should be in the factory - that way it'll get run auto-magically for all the alt-VMs once they upgrade?
There was a problem hiding this comment.
Done — wired in the factory (startAdminConsole in cmd/verifier/adminconsole.go, called from the committee factory), so alt-VMs get it on upgrade. #1501
| // A present console config serves the admin UI as a sibling process in this | ||
| // container; the verifier's own lifecycle is unaffected. | ||
| if stopConsole := cmd.StartAdminConsoleSibling(); stopConsole != nil { | ||
| defer stopConsole() | ||
| } |
There was a problem hiding this comment.
Same — the token factory calls the same startAdminConsole helper. #1501
| path := os.Getenv(admin.ConfigPathEnv) | ||
| if path == "" { | ||
| path = admin.DefaultConfigPath | ||
| } |
There was a problem hiding this comment.
Haven't looked yet, but if the admin console is verifier-only (to start) was wondering if there should be a separate config/secret file for that, vs. embedding it in the existing config files (e.g. verifier config, executor config, etc.)
There was a problem hiding this comment.
Kept a separate (now tiny) config file as the enable signal — it carries only listen_address, optional aggregator_address/trace_url, and access.actor_header. With the single-verifier scoping there are no per-node secrets paths anymore, so the file no longer references anything only the deployer knows; the credential lives in the verifier secrets file's [admin_ui] table. Embedding in the job spec would work too now, but the file keeps the committee/token wiring identical and keeps console settings out of JD.
| `admin-console/config.documented.toml` is the exception to the generation rule below: it is | ||
| hand-written in the same style until a `tools/configdoc` target is registered for | ||
| `verifier/pkg/admin` (that package has fields without doc comments, which the completeness | ||
| gate rejects). Keep it in sync with the struct by hand; replace it with generated output | ||
| once the target exists. |
There was a problem hiding this comment.
Can you add the doc comments and add the target? It shouldn't be a big lift hopefully.
There was a problem hiding this comment.
Done — every field has a doc comment and there's a registered tools/configdoc target, so admin-console/config.documented.toml is generated by just config-docs like the rest (the hand-written exception note is removed from the README). #1502
| # nodes are the verifier databases this console administers. Nodes must belong to the same | ||
| # operator; each entry is one verifier's application database. At least one is required and | ||
| # names must be unique. | ||
| [[nodes]] |
There was a problem hiding this comment.
If we simplify things so that a single admin console only accesses that verifier's DB/state, I think we can simplify things further?
There was a problem hiding this comment.
Done — one console per verifier, sharing that verifier's application DB. The [[nodes]] model, node secrets paths, and the separate console DB/read-only mode are all gone. #1500
| // IndexerURL (optional base URL) enables the indexer's verification-result lookup. | ||
| IndexerURL string `toml:"indexer_url"` |
There was a problem hiding this comment.
I think we can remove this, since we're scoping to verifier-only
There was a problem hiding this comment.
Removed — attestation freshness checks are aggregator-only now.
| // TraceURL (optional) is a base URL to the operator's trace viewer, linked from the | ||
| // message detail page when set. | ||
| TraceURL string `toml:"trace_url"` |
There was a problem hiding this comment.
What is this typically? Like victoria traces or something?
There was a problem hiding this comment.
Typically an internal trace viewer — Grafana/Tempo or Jaeger. The doc comment (and the generated config reference) now says that.
| "github.com/smartcontractkit/chainlink-common/pkg/logger" | ||
| ) | ||
|
|
||
| const gooseTableName = "ccv_admin_goose_db_version" |
There was a problem hiding this comment.
Following up on making this verifier-only for now: do we need a separate set of migrations to manage + a separate DB? Can we put the migrations in the verifier DB instead?
There was a problem hiding this comment.
Done — ccv_admin_actions is created by the verifier's own migrations (verifier/migrations/postgres/00010_admin_actions.sql); the separate console DB, goose table, and migration runner are deleted. #1500
|
Oh one more thing - its a big PR, if there's a way we can break it down a bit more that'd be cool, maybe adding the admin package first, then integrating it w/ the factories, but there might be better splits. |
- one console per verifier, sharing its app DB and secrets file (no [[nodes]], no console DB, no read-only mode) - console served in-process by the factories; drop ccv admin serve and the sibling process - action log migrates with the verifier migrations (00010) - attestation freshness checks are aggregator-only - register the configdoc target for the admin console config
|
Code coverage report:
Files added (in
|
|
Thanks for the review — all of the simplifications landed, and this PR is now split into a stack:
The Copilot findings were either already fixed on head (audit intent ordering, execute-time gating, attestation error mapping) or are gone with the rescope (backfill, dashboard, multi-node DB fallback) — replied to each thread individually. Closing this in favor of the stack. |


Description
https://smartcontract-it.atlassian.net/browse/CCIP-13503
https://smartcontract-it.atlassian.net/browse/CCIP-13505
Adds the CCV admin console: a server-rendered web UI (
templ/htmx) for finding and recovering dropped messages, served from the verifier's own container as a supervised sibling process on a dedicated admin UI port. It wraps the same job-queue and recovery machinery theccv job-queueandccv recoveryCLIs drive — same semantics, same tables — and administers committee and token verifier databases only.A console config at
/etc/ccv-admin/config.toml(override:CCV_ADMIN_CONFIG_PATH) enables it: the committee and token entrypoints spawnccv admin serveas a child and respawn it if it crashes; no config file means the console stays off. The verifier never needs a restart just to administer the console, and a console crash never touches the verifier.Features
GetVerifierResultsForMessagereads, indexer lookup otherwise): attested targets are excluded from execution; an unknown state disables the target rather than proving a replay is needed.Safety model (details in
docs/verifier/admin-console.md)listen_addressrequiresaccess.actor_header(identity from an authenticating proxy) and fails validation otherwise.Content-Security-Policy: default-src 'self',X-Frame-Options: DENY,Referrer-Policy: no-referrer).One migration for the console's own database (
ccv_admin_actions, the action log), applied with an instance-scoped goose provider so it never collides with verifier migrations. The protobuf aggregator client lives inintegration/storageaccess(ResultsClient+ResultEntry) because verifier packages must not import chainlink protos.Standalone verifier only. Chainlink core-node deployments intentionally apply neither the console nor its migration.
Testing
go test ./verifier/pkg/admin/... ./cmd/verifier/... -shuffle on -race— green, including repeated shuffle runs (migration ordering), Docker-backed Postgres round-trips for the action log and migrations, and the sibling-supervisor lifecycle.integration/storageaccesscover the protobuf →ResultEntrytranslation; admin-side tests cover each attestation entry state, the reschedule gate, recovery submit/cancel/resume, and the action-log intent/outcome contract.golangci-lint runclean on touched packages;just tidyidempotent; config-doc freshness test passes.Checklist
changelogdirectory)