Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 52 additions & 0 deletions changelog/2026-09-23_admin_console.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# Admin console for verifier recovery (U1–U3)

## Executive Summary

- Adds a server-rendered admin console (templ/htmx, Gin) shipped in both verifier images and
served in-process by the verifier when a console config file is present, wrapping the
job-queue and recovery stores so operators can find, explain, and recover dropped messages
without node shell access or CLI flags.
- The console administers the verifier it runs beside — one console, one verifier. It shares
that verifier's application database and secrets file, so there is nothing extra to
provision: no console database, no per-node secrets references.
- Covers message search in the verifier's failed-job archive with lookup-failure-vs-empty
separation, a per-message detail page (failure category, archive age/expiry, durable
drop/incident evidence with coverage window, chain-status context), attestation freshness
checks (anonymous aggregator reads) gating reschedule, owner-scoped reschedule with preview
and per-target outcomes, and a durable action log.
- Source-range recovery (replay/reset-reader) is driven through the durable R5 operations with
progress, cancel/resume, and reload-safe tracking; R4 evidence is shown alongside the chosen
range. Indexer-data backfill is out of scope for now (deferred with the indexer admin UI);
indexer repair stays with the indexer's own replay tooling.
- Safety model: loopback bind by default (non-loopback requires an identity source — an
authenticating-proxy actor header or `[admin_ui]` basic auth from the verifier secrets
file), CSRF-protected mutations, and every mutation recorded with an intent row before it
runs — an unaudited mutation never proceeds.
- Console state is one Postgres table (`ccv_admin_actions`) created by the verifier's own
migrations, alongside the stores it administers. No changes to verifier runtime behavior
when the config file is absent.
- Packaging: a console config at `/etc/ccv-admin/config.toml` (`CCV_ADMIN_CONFIG_PATH`)
enables the console; both verifier factories (committee and token, so alt-VMs inherit it)
serve it in-process on its own port and shut it down with the job. No config file means
disabled.

## AI Adapter Index

Purely additive except for the CLI command table. Unlisted symbols keep their existing contracts.

| Symbol | Kind | Search | Location |
| --- | --- | --- | --- |
| `admin` package (console) | added | `verifier/pkg/admin` | `verifier/pkg/admin/` |
| `cli/admin.Command` (`ccv admin check-config`) | added | `admin\.Command` | `cli/admin/commands.go` |
| `startAdminConsole` (factory wiring) | added | `startAdminConsole` | `cmd/verifier/adminconsole.go` |
| `admin.BasicAuthFromSecrets / ValidateAccessPolicy` | added | `BasicAuthFromSecrets` | `verifier/pkg/admin/auth.go` |
| `vsecrets.VerifierSecrets.AdminUIAuth` | added | `AdminUIAuth` | `verifier/pkg/vsecrets/vsecrets.go` |
| `[admin_ui]` secrets table | added | `admin_ui` | `docs/config/verifier/secrets.documented.toml` |
| `ccv_admin_actions` table | added | `00010_admin_actions` | `verifier/migrations/postgres/00010_admin_actions.sql` |

## Compatibility

The console administers the standalone verifier's own application database. It only uses the
live-safe operations; the offline-only `ccv chain-statuses` mutations are deliberately not
exposed. The Chainlink-node integration is untouched: the console is wired in the standalone
factories only.
49 changes: 49 additions & 0 deletions cli/admin/commands.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
// Package admin provides the `ccv admin` commands. The console itself is served
// in-process by the verifier factory when the config file is present; this group is
// for pre-flight validation of that file.
package admin

import (
"fmt"

"github.com/urfave/cli"

"github.com/smartcontractkit/chainlink-ccv/verifier/pkg/admin"
)

// Command returns the `ccv admin` command group.
func Command() cli.Command {
return cli.Command{
Name: "admin",
Usage: "Admin console helpers (the console is served by the verifier process itself)",
Subcommands: []cli.Command{
{
Name: "check-config",
Usage: "Validate the console config file the verifier would load at startup",
Flags: []cli.Flag{
cli.StringFlag{
Name: "config",
Usage: "Path to the console config TOML",
EnvVar: admin.ConfigPathEnv,
Value: admin.DefaultConfigPath,
},
},
Action: func(c *cli.Context) error {
cfg, err := admin.LoadConfig(c.String("config"))
if err != nil {
return err
}
access := "actor local (loopback)"
if cfg.Access.ActorHeader != "" {
access = "proxy header " + cfg.Access.ActorHeader
}
fmt.Println("config OK: listen=" + cfg.ListenAddress + " access=" + access) //nolint:forbidigo // CLI user output
if cfg.AggregatorAddress != "" {
fmt.Println(" attestation freshness checks via aggregator " + cfg.AggregatorAddress) //nolint:forbidigo // CLI user output
}
return nil
},
},
},
}
}
4 changes: 3 additions & 1 deletion cli/recovery/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
# CCV live recovery CLI

The standalone verifier accepts durable recovery requests through its existing PostgreSQL database. The running source reader performs the work on its event loop. There is no admin HTTP endpoint or UI in this change. These commands require a binary and schema containing migration 00009; the existing verifier migration mechanism applies it during upgrade. Chainlink core must separately expose this command group before it is available through `chainlink node`.
The standalone verifier accepts durable recovery requests through its existing PostgreSQL database. The running source reader performs the work on its event loop. These commands require a binary and schema containing migration 00009; the existing verifier migration mechanism applies it during upgrade. Chainlink core must separately expose this command group before it is available through `chainlink node`.

A server-rendered admin console wrapping these flows is served in-process by the standalone verifier when its config file is present; see `docs/verifier/admin-console.md`.

Idle readers check for new recovery operations every 15 seconds (`sourcereader.RecoveryPollInterval`), so a submission can take up to that long to be picked up; an active operation runs at full event-loop speed. The coarse idle cadence keeps the control-plane database reads negligible.

Expand Down
64 changes: 64 additions & 0 deletions cmd/verifier/adminconsole.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
package verifier

import (
"context"
"fmt"
"os"
"sync"

"github.com/jmoiron/sqlx"

"github.com/smartcontractkit/chainlink-ccv/verifier/pkg/admin"
"github.com/smartcontractkit/chainlink-ccv/verifier/pkg/vsecrets"
"github.com/smartcontractkit/chainlink-common/pkg/logger"
"github.com/smartcontractkit/chainlink-common/pkg/sqlutil"
)

// startAdminConsole serves the admin console in-process when its config file is
// present (CCV_ADMIN_CONFIG_PATH or /etc/ccv-admin/config.toml); the console shares
// the verifier's application DB and its [admin_ui] credential. Absent file means
// disabled (nil, nil); the returned stop function shuts the console down.
func startAdminConsole(lggr logger.Logger, ds sqlutil.DataSource, secrets *vsecrets.VerifierSecrets, aggregatorAddress string) (func(), error) {
path := os.Getenv(admin.ConfigPathEnv)
if path == "" {
path = admin.DefaultConfigPath
}
if _, err := os.Stat(path); err != nil { //nolint:gosec // G703: operator-provided config path, not request input.
return nil, nil
}
cfg, err := admin.LoadConfig(path)
if err != nil {
return nil, err
}
db, ok := ds.(*sqlx.DB)
if !ok || db == nil {
return nil, fmt.Errorf("admin console requires the verifier application database ([db].url in the verifier secrets file)")
}
auth, err := admin.BasicAuthFromSecrets(secrets)
if err != nil {
return nil, err
}
if cfg.AggregatorAddress == "" {
cfg.AggregatorAddress = aggregatorAddress
}
srv, err := admin.NewServer(cfg, admin.Deps{DB: db, Auth: auth, AggregatorAddress: cfg.AggregatorAddress}, lggr)
if err != nil {
return nil, err
}

ctx, cancel := context.WithCancel(context.Background())
done := make(chan struct{})
go func() {
defer close(done)
if err := srv.Run(ctx); err != nil {
lggr.Errorw("admin console stopped with error", "error", err)
}
}()
var once sync.Once
return func() {
once.Do(func() {
cancel()
<-done
})
}, nil
}
2 changes: 2 additions & 0 deletions cmd/verifier/run_ccv_cli.go
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ import (
"go.uber.org/zap"
"go.uber.org/zap/zapcore"

"github.com/smartcontractkit/chainlink-ccv/cli/admin"
"github.com/smartcontractkit/chainlink-ccv/cli/chainstatuses"
"github.com/smartcontractkit/chainlink-ccv/cli/jobqueue"
"github.com/smartcontractkit/chainlink-ccv/cli/migrate"
Expand Down Expand Up @@ -115,6 +116,7 @@ func RunCCVCLI(args []string, secretsEnvVar, defaultSecretsPath string) {
Usage: "CCV-related commands",
Subcommands: []cli.Command{
{Name: "recovery", Usage: "Live source-range recovery and durable admission evidence", Subcommands: recoverycli.InitCommandsWithFactory(getRecoveryStore)},
admin.Command(),
{
Name: "chain-statuses",
Usage: "List, enable, disable, or set finalized block height for chain statuses",
Expand Down
19 changes: 19 additions & 0 deletions cmd/verifier/servicefactory.go
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ type factory struct {
aggregatorWriter *storageaccess.FanOutWriter
heartbeatClient heartbeatclient.HeartbeatSender
chainStatusDB sqlutil.DataSource
adminStop func()
}

var _ bootstrap.ServiceFactoryValidator = (*factory)(nil)
Expand Down Expand Up @@ -473,6 +474,18 @@ func (f *factory) Start(ctx context.Context, spec bootstrap.JobSpec, deps bootst
f.server = server
f.coordinator = coordinator

// The admin console serves in-process when its config file is present; it shares
// this verifier's application database and secrets.
aggregatorAddress := ""
if len(resolvedAggregators) > 0 {
aggregatorAddress = resolvedAggregators[0].Address
}
adminStop, err := startAdminConsole(lggr, chainStatusDB, secrets, aggregatorAddress)
if err != nil {
return fmt.Errorf("failed to start admin console: %w", err)
}
f.adminStop = adminStop

lggr.Infow("🎯 Verifier service fully started and ready!")

return nil
Expand Down Expand Up @@ -531,11 +544,17 @@ func (f *factory) Stop(ctx context.Context) error {
}
}

// Stop the admin console
if f.adminStop != nil {
f.adminStop()
}

f.server = nil
f.coordinator = nil
f.profiler = nil
f.aggregatorWriter = nil
f.heartbeatClient = nil
f.adminStop = nil
f.lggr = nil
f.chainStatusDB = nil

Expand Down
13 changes: 13 additions & 0 deletions cmd/verifier/tokenfactory.go
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ type tokenVerifierFactory struct {

coordinators []*verifier.Coordinator
httpServer *http.Server
adminStop func()
lggr logger.Logger
}

Expand All @@ -49,6 +50,10 @@ func NewTokenVerifierServiceFactory() bootstrap.ServiceFactory {
// Stop tries to stop all services gracefully.
func (tvf *tokenVerifierFactory) Stop(_ context.Context) error {
var errs []error
if tvf.adminStop != nil {
tvf.adminStop()
tvf.adminStop = nil
}
if tvf.httpServer != nil {
// Graceful shutdown
shutdownCtx, shutdownCancel := context.WithTimeout(context.Background(), 30*time.Second)
Expand Down Expand Up @@ -131,6 +136,14 @@ func (tvf *tokenVerifierFactory) Start(ctx context.Context, spec bootstrap.JobSp
return fmt.Errorf("failed to connect to Postgres database: %w", err)
}

// The admin console serves in-process when its config file is present; it shares
// this verifier's application database and secrets. The token verifier has no
// aggregator of its own, so freshness checks need aggregator_address in the file.
tvf.adminStop, err = startAdminConsole(tvf.lggr, db, secrets, "")
if err != nil {
return fmt.Errorf("failed to start admin console: %w", err)
}

postgresStorage := storage.NewPostgres(db, tvf.lggr)
// Wrap storage with monitoring decorator to track query durations
monitoredStorage := storage.NewMonitoredStorage(postgresStorage, verifierMonitoring.Metrics())
Expand Down
1 change: 1 addition & 0 deletions docs/config/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ This directory holds the config and secrets reference for every CCV app, one
| indexer | `indexer/config.documented.toml`, `indexer/secrets.documented.toml` |
| bootstrap | `bootstrap/config.documented.toml`, `bootstrap/secrets.documented.toml` |
| monitoring (shared) | `common/monitoring.documented.toml` |
| admin console | `admin-console/config.documented.toml` |

Each file is a working TOML document: the values are the app's defaults where a
default exists, and illustrative examples otherwise, and every field is annotated
Expand Down
23 changes: 23 additions & 0 deletions docs/config/admin-console/config.documented.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
# Code generated by tools/configdoc. DO NOT EDIT.
# Admin console configuration reference. Values shown are defaults or illustrative examples.

# listen_address is the bind address; loopback by default.
listen_address = "127.0.0.1:8105"

# aggregator_address (optional, host:port) overrides the aggregator used for
# attestation freshness checks via the unauthenticated GetVerifierResultsForMessage.
# Empty uses the verifier's own first configured aggregator.
aggregator_address = "aggregator-1:50051"

# trace_url (optional) is a base URL to the operator's trace viewer — typically an
# internal Grafana/Tempo or Jaeger — linked from the message detail page when set.
trace_url = "https://traces.example.com"

# access configures how the console identifies who is acting.
[access]
# actor_header names the HTTP header carrying an authenticated identity from a
# fronting proxy (shared hosting). Empty means self-hosted loopback: actor "local".
# Non-loopback serving requires this header or [admin_ui] basic auth from the
# verifier secrets file (validated at startup, when the secrets are loaded).
actor_header = "X-Authenticated-User"

9 changes: 9 additions & 0 deletions docs/config/verifier/secrets.documented.toml
Original file line number Diff line number Diff line change
Expand Up @@ -21,3 +21,12 @@
# secret_key is the HMAC secret the request signature is computed with.
secret_key = "<secret-key>"

# admin_ui is the optional basic-auth credential gating the admin console UI. Only the
# admin console consumes it, when this file is the console's secrets file; the verifier
# binaries ignore it.
[admin_ui]
# username is the basic-auth username; it also becomes the action-log actor.
username = "operator"
# password is the basic-auth password.
password = "<password>"

14 changes: 11 additions & 3 deletions docs/runbooks/remediating-stuck-or-dropped-messages.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,10 @@
# Runbook: Remediating a Stuck or Dropped Message

_Last reviewed: 2026-09-10._
_Last reviewed: 2026-09-23._

Use after [unverified-message triage](./unverified-message-after-15-minutes.md) or [unexecuted-message triage](./unexecuted-message-after-15-minutes.md) identifies the affected owner, source and messages. Recovery is per affected committee member and database. Cross-node discovery/fan-out remains an operator or deployment-layer responsibility.
Use after [unverified-message triage](./unverified-message-after-15-minutes.md) or [unexecuted-message triage](./unexecuted-message-after-15-minutes.md) identifies the affected owner, source and messages. Recovery is per affected committee member and database.

When the [admin console](../verifier/admin-console.md) is deployed, it is the primary path: each verifier serves its own console in-process, driving every action below from the browser and recording each mutation in its action log. The console administers the one verifier it runs beside, so repeat the flow per affected committee member; the CLI steps in this runbook remain the documented fallback.

## 1. Pick the Lever

Expand All @@ -17,6 +19,8 @@ Use after [unverified-message triage](./unverified-message-after-15-minutes.md)

**Reschedule uses the saved payload and skips source-reader finality, curse and disablement admission checks.** It is unsuitable for deciding whether an event remains canonical after a reorg. Source recovery re-reads events that still exist on the chain and enters ordinary verification/policy processing after admission. Neither path bypasses policy. Indexer backfill refreshes the indexer's view of results; it does not re-admit verifier source events or retry policy decisions.

In plain language: use a **reschedule** when the verifier already holds the message — a failed job retained in its archive — and the fix is to run verification and policy (or just persistence) again on the saved payload. Use **source replay** when the verifier never admitted the message (curse/rule drop, missed interval, expired archive) and the source chain must be re-read to decide. Use the **investigated reader reset** for the replay special case of a finality-disabled reader, and **indexer backfill** when the verifier and aggregator are fine and only the indexer's view needs repair (the indexer's own replay tooling; not a console action). [The admin console guide](../verifier/admin-console.md#the-recovery-actions) walks through what each action does and does not do; in the console these are the actions on the message detail and source recovery pages rather than CLI invocations.

## 2. Check the Time Windows

Automatic retry remains **7 days**, with non-retryable failures (including policy FAIL) archived immediately. Archive retention remains **30 days after archiving**, swept every 4 hours. The message's creation time does not start that retention window.
Expand All @@ -40,6 +44,8 @@ Drop evidence is separate from archives. It is retained for 30 days since its la

## 3. Reschedule a Single Dropped Message

**Console path:** search the full message ID on the console's message search page, open the message detail, and use the reschedule action there. The preview shows the exact owners and jobs the reschedule will touch and rechecks attestation state before anything mutates; the action is recorded in the console's action log. The CLI steps below are the fallback.

1. Resolve the cause first. A policy endpoint must return PASS for the message before replay can succeed. Confirm that the source event remains valid and the message has not already been attested through another path.
2. Point the CLI at the affected member's database and find the full message IDs:

Expand All @@ -65,6 +71,8 @@ See the [job-queue command reference](../../cli/jobqueue/README.md) and [policy

## 4. Recover a Source Range

**Console path:** the console's source recovery page queries the same drop/incident evidence and submits an ordinary replay or an investigated reader reset with the actor filled from your session and an evidence note required. Operations are durable, so progress, cancel and resume survive page reloads and console restarts. The CLI steps below are the fallback and remain the reference for exact semantics.

### Establish the scope

Identify each affected owner/node and source chain, then query retained evidence:
Expand Down Expand Up @@ -144,4 +152,4 @@ Use [aggregator message-disablement rules](../../aggregator/cli/messagedisableme

## 6. Deployment and Coverage Limits

The new recovery/job-queue commands are exposed by the standalone verifier. Wiring them into Chainlink core, cross-node fan-out, indexer engine changes and an admin UI are outside this change. Owner inference is local to one selected archive queue/database; source recovery always requires an explicit owner. There is no per-message policy bypass. Keep canonical-chain investigation and final-result verification in the operator workflow.
The new recovery/job-queue commands are exposed by the standalone verifier. Wiring them into Chainlink core and indexer engine changes are outside this change; the [admin console](../verifier/admin-console.md) now provides the UI over these flows for the verifier it runs beside. Owner inference is local to one selected archive queue/database; source recovery always requires an explicit owner. There is no per-message policy bypass. Keep canonical-chain investigation and final-result verification in the operator workflow.
Loading
Loading