Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions grafana-alertcheck/.changeset/v0.1.9.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
- Add support for Grafana's `Recovering` instance state ("keep firing for"): a first-seen recovering instance is classified preexisting, and `check` observes it through its own recovery deadline so a genuine recovery reads `recovered` while an unresolved, paused, or absent one stays `still_failing`.
- Remove `watch --poll-interval`. Every rule now polls at exactly half its evaluation interval, so a state that lasts a full interval can never fall between polls; the flag could only widen `maxGap` and hide transitions.
2 changes: 1 addition & 1 deletion grafana-alertcheck/cmd/check.go
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ func runCheck(args []string, stdin io.Reader, stdout, stderr io.Writer) int {
pidfile := fs.String("pidfile", "", "pidfile of the recorder to stop before reading --in (default <in>.pid)")
from := fs.String("from", "", "the moment the deploy finished, RFC3339 (required with --in)")
to := fs.String("to", "", "the end of the window to classify, RFC3339 (required)")
states := fs.String("states", "", "comma-separated bad states to classify against (default: firing)")
states := fs.String("states", "", "comma-separated bad states to classify against (default: firing,recovering)")
preexisting := fs.String("preexisting", "", "how to judge an instance already bad at `from` (default: fail-unless-recovered)")
minObserved := fs.Int("min-observed", 0, "minimum rules that must be observed (default: every resolved rule)")
allowPaused := fs.Bool("allow-paused", false, "do not count a rule paused before the window against --min-observed")
Expand Down
22 changes: 10 additions & 12 deletions grafana-alertcheck/cmd/common.go
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,9 @@ import (
)

// commonFlags is registerCommon's result: the flags watch and check share.
// Connection details are never flags, and states / poll-interval are
// deliberately NOT here — states is check-only because recording is
// unfiltered, and poll-interval is watch-only because check reads the cadence
// from the log header. Putting either here would give both commands an opinion
// about a value only one of them may set.
// Connection details are never flags, and states is deliberately NOT here:
// states is check-only because recording is unfiltered, and putting it here
// would give both commands an opinion about a value only one of them may set.
type commonFlags struct {
folder *string
concurrency *int
Expand Down Expand Up @@ -101,13 +99,13 @@ func readAlerts(stdin io.Reader, flagName, path string) ([]string, error) {
// parseStates parses check's --states flag: a comma-separated list of the
// "bad" state vocabulary Config.States matches against (classify.go's
// badStateSet). An empty string is not resolved here — it means "use the
// library default of {firing}" — so this returns nil, nil for "" rather than
// an error.
// library default of {firing, recovering}" — so this returns nil, nil for ""
// rather than an error.
//
// normal is deliberately NOT accepted. The vocabulary is fixed to
// firing | pending | nodata | error precisely because "normal" is the good
// state, never a bad one to classify against: --states normal would turn every
// healthy instance into a violation and fail every healthy fleet.
// firing | pending | recovering | nodata | error precisely because "normal" is
// the good state, never a bad one to classify against: --states normal would
// turn every healthy instance into a violation and fail every healthy fleet.
func parseStates(s string) ([]gate.State, error) {
if strings.TrimSpace(s) == "" {
return nil, nil
Expand All @@ -119,10 +117,10 @@ func parseStates(s string) ([]gate.State, error) {
continue
}
switch gate.State(part) {
case gate.StateFiring, gate.StatePending, gate.StateNodata, gate.StateError:
case gate.StateFiring, gate.StatePending, gate.StateRecovering, gate.StateNodata, gate.StateError:
out = append(out, gate.State(part))
default:
return nil, fmt.Errorf("--states: unknown state %q (want any of: firing, pending, nodata, error)", part)
return nil, fmt.Errorf("--states: unknown state %q (want any of: firing, pending, recovering, nodata, error)", part)
}
}
if len(out) == 0 {
Expand Down
11 changes: 1 addition & 10 deletions grafana-alertcheck/cmd/watch.go
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ import (

const watchUsage = "usage: grafana-alertcheck watch --out <file> [--pidfile F] [--daemon-log F] " +
"(--alerts <file|-> [--folder F] | --include-labels k=v,... [--exclude-labels k=v,...]) [--exclude-alerts <file|->] " +
"[--poll-interval D] [--concurrency N] [--until RFC3339]"
"[--concurrency N] [--until RFC3339]"

// runWatch is the record step's entire CLI surface, split in two by one flag
// set — gate.DaemonChildFlag ("--daemon-child") and gate.ReadyFDFlag
Expand Down Expand Up @@ -44,7 +44,6 @@ func runWatch(args []string, stdin io.Reader, stdout, stderr io.Writer) int {
pidfile := fs.String("pidfile", "", "pidfile path (default <out>.pid)")
daemonLog := fs.String("daemon-log", "", "stdout/stderr sink for the detached recorder (default <out>.daemon.log)")
until := fs.String("until", "", "optional hard stop, RFC3339 (default: run until check stops it)")
pollInterval := fs.String("poll-interval", "", "override every rule's poll cadence (default: half its own evaluation interval)")

// Hidden: never in watchUsage, never typed by an operator (see doc comment).
daemonChild := fs.Bool(gate.DaemonChildFlag[2:], false, "")
Expand Down Expand Up @@ -124,14 +123,6 @@ func runWatch(args []string, stdin io.Reader, stdout, stderr io.Writer) int {
}
cfg.Until = t
}
if *pollInterval != "" {
d, err := time.ParseDuration(*pollInterval)
if err != nil {
fmt.Fprintf(stderr, "--poll-interval: %v\n", err)
return 2
}
cfg.PollEvery = d
}

if err := gate.Watch(context.Background(), cfg); err != nil {
fmt.Fprintln(stderr, err)
Expand Down
3 changes: 0 additions & 3 deletions grafana-alertcheck/cmd/watch_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -32,9 +32,6 @@ func TestRunWatch_FlagValidation(t *testing.T) {
{"until in the past", true, func(t *testing.T) []string {
return []string{"--out", t.TempDir() + "/log.jsonl", "--alerts", writeTempAlerts(t), "--until", "2000-01-01T00:00:00Z"}
}, "not in the future"},
{"bad poll-interval", true, func(t *testing.T) []string {
return []string{"--out", t.TempDir() + "/log.jsonl", "--alerts", writeTempAlerts(t), "--poll-interval", "not-a-duration"}
}, "--poll-interval"},
{"alerts and labels", true, func(t *testing.T) []string {
return []string{"--out", t.TempDir() + "/log.jsonl", "--alerts", writeTempAlerts(t), "--include-labels", "team=bcm"}
}, "cannot be combined with label selection"},
Expand Down
4 changes: 2 additions & 2 deletions grafana-alertcheck/docs/advanced.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ description: Why grafana-alertcheck schedules per rule, how the request budget w

## Per-rule schedules, never a global cycle

Each rule polls at its **own** cadence (default: half the rule's own evaluation interval). There is deliberately no single global minimum-interval cycle. Overwrite with `--poll-interval`.
Each rule polls at its **own** cadence: half the rule's own evaluation interval, always. The cadence is not configurable — polling faster cannot reveal more (Grafana state only changes on evaluations) and polling slower could let a state fall between polls.

One rule at `intervalSeconds=10` beside twenty at `300` keeps a 5 s cadence for itself and 150 s for the other twenty — not a 5 s cycle for all of them, which would be a 60× request bloat at ~1.8 s per request and would fail to start on a reasonable fleet.

Expand All @@ -25,7 +25,7 @@ The gate records one observation of every rule up front and checks the schedule
- **Burst bound** — the slowest request exceeds the fleet's tightest cadence, which can open a mid-run gap.
- **Startup handoff** — draining the first-observation pass's backlog at `--concurrency` would leave some rule unpolled past its own `maxGap`. A rule the pass observed early is seeded overdue, and a tight rule observed late can queue behind every rule due before it. The gate simulates the poller's first cycles from the recorded observation times and measured latencies — each wake takes every rule due at that instant, polls the batch at `--concurrency`, and wakes again when it ends — and refuses if any rule's first poll would land past its `maxGap`. Steady-state utilization cannot see this — a long pass at low concurrency is exactly the case it passes.

The error names only the levers that can fix it: the minimum `--concurrency` when the schedule is concurrency-bound, and `--poll-interval` or a smaller alert set for single-request shapes concurrency cannot shorten. It never prescribes a single interval.
The error names only the levers that can fix it: the minimum `--concurrency` when the schedule is concurrency-bound, and a smaller alert set for single-request shapes concurrency cannot shorten.

## The startup pass and `ready_at`

Expand Down
11 changes: 7 additions & 4 deletions grafana-alertcheck/docs/how-alerts-are-evaluated.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,12 +19,13 @@ Grafana reports instance states in two vocabularies (`Alerting`/`Normal` at inst
| `normal` | Healthy |
| `firing` | The condition is true and `for` has elapsed |
| `pending` | The condition is true, `for` has not elapsed |
| `recovering` | The condition has cleared but the rule's *keep firing for* has not elapsed |
| `nodata` | The query returned no series (synthetic instance) |
| `error` | The query failed (synthetic instance) |

A rule's **rule-level** `state` and `health` are kept verbatim and only reported — they are never classified. The **instance** state is what the classifier reasons about.

A "bad" instance is one whose canonical state is in `--states` (default `firing`). `pending` and `nodata` are excluded by default.
A "bad" instance is one whose canonical state is in `--states` (default `firing,recovering`). `pending` and `nodata` are excluded by default. `recovering` is bad because the instance is still firing (Alertmanager keeps notifying) until its recovery period ends; `--states firing` opts out of tracking it.

## Verdict model

Expand All @@ -35,7 +36,7 @@ For each instance the gate builds a timeline of bad spans over `[from, to]`, the
| `healthy` | Good throughout, observed throughout | → 0 |
| `new_failure` | Entered a bad state **inside** the window | → 1 |
| `still_failing` | Bad at `from`, still bad at `to` | → 1 |
| `recovered` | Bad at `from`, cleared before `to`, stayed clear | → 0 |
| `recovered` | Bad at `from`, cleared and stayed clear (possibly during the recovery observation) | → 0 |
| `unstable` | Cleared, then became bad again | → 1 |
| `paused` | Paused **before** the window opened | counts against `--min-observed` unless `--allow-paused` |
| `not_verified` | The window could not be observed: a gap, sustained `health=error`, a stale evaluation, or an absent rule | → 2 |
Expand Down Expand Up @@ -91,8 +92,10 @@ A preexisting bad instance is deliberately **not** terminal: if it clears before

Fail-fast is on by default and always preserves the failure: an early run can exit `1` or `2`, never `0`. The one difference from a full run is that an early exit may report `1` before an inability surfaces that would have made it `2`. `--fail-fast=false` disables the guard and always waits for the full window and its coverage proof.

## The drain wait and `transitionGrace`
## The drain wait, `transitionGrace`, and recovery observation

A condition that arises just before `to` becomes `firing` only at the first evaluation after its `for` elapses. `transitionGrace` (derived from the watched rules' `for` values) extends the classification bound past `to` so such a surfacing condition is caught. After collection, a **drain wait** polls until each rule has evaluated through `to + transitionGrace` (bounded by `drainTimeout`); a rule that never does is `not_verified`.

Run time = `(to − from) + transitionGrace + drainTimeout`. This is printed at start. A requested window with a subsecond part is rounded up to the next whole second by extending `to`, so the plan never reads a window like `9m59.99445781s`.
The mirror case is recovery: an instance whose condition cleared just before `to` enters `recovering` and only resolves after its **keep firing for** elapses, which can be long after `to`. A first-seen `recovering` instance is always treated as preexisting — recovering is only reachable from `firing`, and every rule is sampled twice per evaluation interval, so an in-window fire cannot be missed between polls; its `activeAt` is the recovery onset, never the fire onset. The episode stays open until the instance reports `normal`. To decide that, `check` keeps observing the affected rules past `to + transitionGrace`, up to `activeAt + keep firing for` plus an evaluation/cadence margin. In recorder mode the recorder keeps polling so the extension evidence lands in the log; in single-step mode `check` polls the affected rules directly. Either way the extension is scoped to the recovering instances: a different instance going bad during the observation is outside the window and stays `healthy`. Each instance expires on its **own** deadline, and a clear past it — or one arriving after the rule paused or disappeared — does not resolve the episode: it stays `still_failing` (fail-closed).

Run time = `(to − from) + transitionGrace + drainTimeout`, plus the recovery observation when one triggers; the extra deadline is printed when it starts. A requested window with a subsecond part is rounded up to the next whole second by extending `to`, so the plan never reads a window like `9m59.99445781s`.
5 changes: 2 additions & 3 deletions grafana-alertcheck/docs/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ grafana-alertcheck list
```bash
grafana-alertcheck watch --out <file> [--pidfile F] [--daemon-log F] \
(--alerts <file|-> [--folder F] | --include-labels k=v,... [--exclude-labels k=v,...]) \
[--exclude-alerts <file|->] [--poll-interval D] [--concurrency N] [--until RFC3339]
[--exclude-alerts <file|->] [--concurrency N] [--until RFC3339]
```

| Flag | Default | Meaning |
Expand All @@ -40,7 +40,6 @@ grafana-alertcheck watch --out <file> [--pidfile F] [--daemon-log F] \
| `--include-labels` | — | Comma-separated exact-match `key=value` pairs selecting rules by label (cannot be combined with `--alerts`) |
| `--exclude-labels` | — | Comma-separated exact-match `key=value` pairs; a rule carrying any of them is dropped (requires `--include-labels`) |
| `--exclude-alerts` | — | File of alert names, one per line, or `-` for stdin; subtracted from the selected set (works with `--alerts` and with labels) |
| `--poll-interval` | half the rule's interval | Override every rule's cadence (never clamped) |
| `--concurrency` | `1` | Max concurrent requests to Grafana |
| `--until` | run until signalled | Optional hard stop |

Expand Down Expand Up @@ -83,7 +82,7 @@ grafana-alertcheck check [--in <file>] [--pidfile F] --from RFC3339 --to RFC3339
| `--include-labels` | — | Comma-separated exact-match `key=value` pairs selecting rules by label (cannot be combined with `--alerts`) |
| `--exclude-labels` | — | Comma-separated exact-match `key=value` pairs; a rule carrying any of them is dropped (requires `--include-labels`) |
| `--exclude-alerts` | — | File of alert names, one per line, or `-` for stdin; subtracted from the selected set (works with `--alerts` and with labels; refused **with** `--in`) |
| `--states` | `firing` | Comma-separated bad states: `firing,pending,nodata,error` |
| `--states` | `firing,recovering` | Comma-separated bad states: `firing,pending,recovering,nodata,error` |
| `--preexisting` | `fail-unless-recovered` | `fail-unless-recovered` \| `fail` \| `ignore` |
| `--min-observed` | every resolved rule | Minimum rules that must be observed |
| `--allow-paused` | `false` | Don't count pre-window-paused rules against `--min-observed` |
Expand Down
2 changes: 1 addition & 1 deletion grafana-alertcheck/docs/reference/log-format.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ The header must be line 1, appear once, and carry `schema_version` `1` (any othe
- `url` and `rules` are the log's identity — `check` validates them against the current environment and a fresh ruler read.
- `started_at` is when the recording opened; `ready_at` is when the first-observation pass completed and every watched, non-paused rule had been observed once. The pass is sequential, so `check` refuses a `from` before `ready_at` (a window opening inside the pass would rest on observations that do not exist). `ready_at` is absent on logs written before the field existed; `check` then falls back to `started_at`.
- `is_paused` records the pause state at record start (the moment `paused` means).
- `poll_every_seconds` is the cadence the recording **actually used** (after any `--poll-interval` override). `check` derives `maxGap` from it, never from `interval_seconds`.
- `poll_every_seconds` is the cadence the recording used: always half the rule's `interval_seconds`. `check` derives `maxGap` from it, never by re-deriving from `interval_seconds`.
- `for_seconds`, `interval_seconds`, `no_data_state`, `exec_err_state` are forensic only — `check` re-resolves definitions and never reads them back.

## Poll
Expand Down
Loading
Loading