Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions grafana-alertcheck/.changeset/v0.1.10.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
- `list`, `watch` and `check` can now observe **datasource-managed** (Prometheus-flavored) alerting rules, auto-discovered from `/api/datasources` (strict `type == "prometheus"` and `jsonData.manageAlerts == true`, then probed) and read through Grafana's per-datasource Prometheus API. No new selection flag: the watch set is still `--alerts`/labels. Loki is deferred.
- `list` gains `DATASOURCE` and `KEY` columns. Datasource rules accept `Group/Title`, `DatasourceName/Group/Title`, and exact `key:` names; `--folder` stays Grafana-only, and ambiguity errors name the datasource.
- Datasource-managed identity is a separate rule key; `uid` stays empty for these rules. The log gains additive `key`, `source_kind`, `datasource_uid`, `datasource_name`, `file` and `rule_key` fields, with `schema_version` still `1`. An old v1 log without `rule_key` remains readable; a new log read by an old binary fails closed on the unresolved key.
- Datasource recovery semantics: because the Prometheus API returns only active instances, an instance leaving the active set is treated as a real recovery (`cleared`). Documented weaker guarantee: a vanished series is indistinguishable from a resolution.
- Pause is not observable for datasource-managed rules (no `isPaused` signal), so check 7 is skipped with an explicit note. `health: err` normalizes to `error` and still triggers check 4.
- Datasource-managed rules record the backend's keep-firing-for (`keep_firing_for` on vmalert, `keepFiringFor` on Prometheus/Mimir; `keep_firing_for_ms` in the log). The datasource API has no `recovering` state — it keeps an alert firing through its keep-firing-for and then drops it — so the recovery observation applies to Grafana-managed rules only. A datasource rule with an unrecognized `type` is now a hard error rather than silently dropped.
- The token now needs `datasources:read` plus datasource query permission even for Grafana-only runs; discovery failures name the missing permission.
- A backend can serve two distinct datasource-managed rules under one identity (same datasource/group/name/file, differing by labels/query). Loading and `list` accept them, but a selection that includes such a rule fails closed: the state query cannot tell the siblings apart, so narrowing to one does not make it observable. `uid:` selectors no longer trigger a bulk datasource read, and a datasource rule's `lastError` is recorded.
- A datasource-managed rule's name can itself contain `/` (e.g. `devex-cicd/prod/griddle-github: ContainersNotReady`). The full name is now matched exactly and used as the `rule_name[]` fetch filter, instead of splitting it into `Group/Title` segments and filtering by the last segment.
- The `RESULTS` table gains a `SOURCE` column (`grafana`/`datasource`). Datasource caveats are printed once before the table, and `DETAILS` carries only rule-specific notes without repeating the alert name. `--output json` adds `source_kind` and a run-level `caveats`.
- Recording rules are never observed: a datasource rules response drops them at parse time, so one can no longer shadow a same-named alerting rule in state selection.
- A no-match error no longer lists substring suggestions; it points at `list`.
26 changes: 23 additions & 3 deletions grafana-alertcheck/cmd/list.go
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,8 @@ func runList(args []string, stdout, stderr io.Writer) int {
return 2
}

defs, err := src.Definitions(context.Background())
fmt.Fprintln(stderr, "discovering rule sources and reading definitions (this can take seconds per source)...")
defs, err := gate.ListAllDefinitions(context.Background(), src)
if err != nil {
fmt.Fprintf(stderr, "reading rule definitions: %v\n", err)
return 2
Expand All @@ -62,9 +63,11 @@ func runList(args []string, stdout, stderr io.Writer) int {
})

tw := tabwriter.NewWriter(stdout, 0, 4, 2, ' ', 0)
fmt.Fprintln(tw, "KIND\tFOLDER\tGROUP\tTITLE\tUID")
fmt.Fprintln(tw, "KIND\tDATASOURCE\tFOLDER\tGROUP\tTITLE\tKEY\tUID")
for _, d := range defs {
fmt.Fprintf(tw, "%s\t%s\t%s\t%s\t%s\n", kindLabel(d.Kind), d.Folder, d.Group, d.Title, uidOrDash(d.UID))
fmt.Fprintf(tw, "%s\t%s\t%s\t%s\t%s\t%s\t%s\n",
kindLabel(d.Kind), datasourceOrDash(d.DatasourceName), d.Folder, d.Group, d.Title,
keyOrDash(d.Key, d.UID), uidOrDash(d.UID))
}
if err := tw.Flush(); err != nil {
fmt.Fprintf(stderr, "writing output: %v\n", err)
Expand All @@ -90,3 +93,20 @@ func uidOrDash(uid string) string {
}
return uid
}

func datasourceOrDash(name string) string {
if name == "" {
return "-"
}
return name
}

func keyOrDash(key, uid string) string {
if key == "" {
key = uid
}
if key == "" {
return "-"
}
return key
}
37 changes: 37 additions & 0 deletions grafana-alertcheck/cmd/list_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,8 @@ func grafanaTestServer(t *testing.T, version string) *httptest.Server {
_, _ = w.Write([]byte(healthBody(version)))
case "/api/ruler/grafana/api/v1/rules":
_, _ = w.Write([]byte(rulerBody))
case "/api/datasources":
_, _ = w.Write([]byte(`[]`))
default:
t.Errorf("unexpected path %q", r.URL.Path)
w.WriteHeader(http.StatusNotFound)
Expand All @@ -68,6 +70,41 @@ func TestRunList_HappyPath(t *testing.T) {
require.Contains(t, out, "grafana-managed")
}

func TestRunList_IncludesDatasourceRules(t *testing.T) {
const dsBody = `{"status":"success","data":{"groups":[{"name":"ExampleMetrics","file":"/etc/vm/rules/example.yml","interval":60,"rules":[
{"name":"ExampleTargetDown","type":"alerting","health":"ok","state":"firing","query":"up == 0","duration":300,
"labels":{"severity":"warning"},"alerts":[]}]}]}}`
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
switch r.URL.Path {
case "/api/health":
_, _ = w.Write([]byte(healthBody("13.1.0")))
case "/api/ruler/grafana/api/v1/rules":
_, _ = w.Write([]byte(rulerBody))
case "/api/datasources":
_, _ = w.Write([]byte(`[{"uid":"vm","name":"VM Prod","type":"prometheus","jsonData":{"manageAlerts":true}}]`))
case "/api/prometheus/vm/api/v1/rules":
_, _ = w.Write([]byte(dsBody))
default:
t.Errorf("unexpected path %q", r.URL.Path)
w.WriteHeader(http.StatusNotFound)
}
}))
t.Cleanup(srv.Close)
t.Setenv("GRAFANA_URL", srv.URL)
t.Setenv("GRAFANA_TOKEN", "test-token")

var stdout, stderr bytes.Buffer
code := run([]string{"list"}, &stdout, &stderr)
require.Equal(t, 0, code)
out := stdout.String()
require.Contains(t, out, "DATASOURCE")
require.Contains(t, out, "KEY")
require.Contains(t, out, "VM Prod")
require.Contains(t, out, "ExampleTargetDown")
require.Contains(t, out, "ds:")
}

func TestRunList_UnsupportedVersion(t *testing.T) {
srv := grafanaTestServer(t, "12.5.0")
t.Setenv("GRAFANA_URL", srv.URL)
Expand Down
40 changes: 34 additions & 6 deletions grafana-alertcheck/cmd/table.go
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ import (
"io"
"sort"
"strconv"
"strings"
"text/tabwriter"
"time"

Expand Down Expand Up @@ -38,7 +39,7 @@ no evaluation for — how long Grafana may go without evaluating the alert befor
func renderTable(w io.Writer, res gate.Result) error {
alertOf := make(map[string]string, len(res.Verdicts))
for _, v := range res.Verdicts {
alertOf[v.RuleUID] = v.Alert
alertOf[verdictKey(v)] = v.Alert
}

enabled := colorEnabled(w)
Expand All @@ -60,11 +61,11 @@ func renderTable(w io.Writer, res gate.Result) error {

fmt.Fprintln(w, "RESULTS")
tw := tabwriter.NewWriter(w, 0, 4, 2, ' ', 0)
fmt.Fprintln(tw, "ALERT\tVERDICT\tBROKEN FOR\tCHECKED EVERY\tWINDOW COVERED\tDETAILS")
fmt.Fprintln(tw, "ALERT\tVERDICT\tBROKEN FOR\tCHECKED EVERY\tWINDOW COVERED\tSOURCE\tDETAILS")
for _, v := range sortedVerdicts(res.Verdicts) {
fmt.Fprintf(tw, "%s\t%s\t%s\t%s\t%s\t%s\n",
fmt.Fprintf(tw, "%s\t%s\t%s\t%s\t%s\t%s\t%s\n",
v.Alert, v.Outcome, v.BadFor.Round(time.Second), v.PollEvery.Round(time.Second),
provedLabel(res.Coverage[v.RuleUID]), v.Note)
provedLabel(res.Coverage[verdictKey(v)]), v.SourceKind, details(v.Alert, v.Note))
}
if err := tw.Flush(); err != nil {
return fmt.Errorf("render table: %w", err)
Expand Down Expand Up @@ -133,6 +134,16 @@ func violationsLabel(n int, enabled bool) string {
return mark + " " + s
}

// details strips the redundant `rule "<title>": ` prefix every coverage note
// carries for the JSON consumer — the ALERT column already names the rule, and
// keeping it would repeat a long name in every DETAILS cell.
func details(title, note string) string {
if note == "" {
return ""
}
return strings.ReplaceAll(note, fmt.Sprintf("rule %q: ", title), "")
}

// provedLabel is the table's WINDOW COVERED column: "yes" for a fully
// observed window, "no" with the reason and largest gap for a not-verified
// rule, and "-" for a rule decide never asked proveCoverage about at all
Expand Down Expand Up @@ -160,12 +171,29 @@ func alertLabel(v gate.Violation, alertOf map[string]string) string {
if v.Alert != "" {
return v.Alert
}
if a, ok := alertOf[v.RuleUID]; ok {
if a, ok := alertOf[violationKey(v)]; ok {
return a
}
return "-"
}

// verdictKey is a RuleVerdict's identity: RuleKey when set, else RuleUID (a
// Result built directly in a test may carry only RuleUID).
func verdictKey(v gate.RuleVerdict) string {
if v.RuleKey != "" {
return v.RuleKey
}
return v.RuleUID
}

// violationKey mirrors verdictKey for a Violation.
func violationKey(v gate.Violation) string {
if v.RuleKey != "" {
return v.RuleKey
}
return v.RuleUID
}

func alertOr(uid string, alertOf map[string]string) string {
if a, ok := alertOf[uid]; ok {
return a
Expand Down Expand Up @@ -201,7 +229,7 @@ func groupedViolations(in []gate.Violation) []violationGroup {
}

func violationSignature(v gate.Violation) string {
return v.Alert + "\x00" + v.RuleUID + "\x00" + string(v.Outcome) + "\x00" + string(v.State) + "\x00" + v.Health + "\x00" + v.Note
return v.Alert + "\x00" + violationKey(v) + "\x00" + string(v.Outcome) + "\x00" + string(v.State) + "\x00" + v.Health + "\x00" + v.Note
}

func sameRendered(a, b gate.Violation) bool {
Expand Down
9 changes: 9 additions & 0 deletions grafana-alertcheck/cmd/table_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,7 @@ func TestRenderTable(t *testing.T) {
require.Contains(t, out, "Zebra Alert")
require.Contains(t, out, "healthy")
require.Contains(t, out, "WINDOW COVERED")
require.Contains(t, out, "SOURCE")

// The violations section must show up even without --output json, and must
// carry the --allow-paused hint text verbatim. INSTANCES is a single word
Expand Down Expand Up @@ -92,6 +93,14 @@ func TestRenderTable(t *testing.T) {
require.Contains(t, out, "❌ violations: 2")
}

// The DETAILS cell drops the `rule "<title>": ` prefix the JSON notes carry, so
// the ALERT column is not repeated in every cell.
func TestDetails_StripsRulePrefix(t *testing.T) {
require.Equal(t, "", details("A", ""))
require.Equal(t, "gap of 5m0s", details("A", `rule "A": gap of 5m0s`))
require.Equal(t, "gap; health=error", details("A", `rule "A": gap; rule "A": health=error`))
}

// The footer verdict line picks its emoji by whether there are violations.
func TestViolationsLabel(t *testing.T) {
require.Equal(t, "✅ violations: 0", violationsLabel(0, false))
Expand Down
10 changes: 10 additions & 0 deletions grafana-alertcheck/docs/advanced.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,16 @@ The detached recorder then continues the schedule the first observations were on

Single-step `check` runs the same pass itself. It cannot watch before it started, so a `from` inside the pass is a declared blind interval: the run warns, classifies from the pass completion, and the live poller continues the pass's schedule. It never classifies a window that opens before every rule has been observed.

## Datasource-managed rules: discovery and cost

Datasource-managed rules are auto-discovered — there is no selection flag. The candidate filter is strict: `type == "prometheus"` **and** `jsonData.manageAlerts == true`. The strict `true` matters: the `AlertStateHistoryBackend` datasource points at the same VictoriaMetrics backend but does not set `manageAlerts: true`, so a `!= false` filter would include it and make every rule name ambiguous. Loki is deferred: its ruler API is broken/disabled in our Grafana, so only Prometheus-flavored rules are in scope.

Each candidate is probed with a `rule_name[]=__probe__` request; a `manageAlerts=true` source whose probe fails is a hard error naming the source, never a silently dropped source.

vmalert allows two distinct alerting rules to share a name within one group. They get the same identity, and the state query is by datasource/group/name/file, so the tool cannot tell them apart: loading and `list` still show both, but a selection that includes either one fails closed — narrowing by a distinguishing label does not help, because the poll would still reduce whichever sibling the backend lists first.

Cost differs by mode. `list` and label selection take the **bulk** response — one request per datasource, several MB and several seconds for a large ruler. Name selection takes a **filtered** request, ~1 KB and ~1 s. The filter uses vmalert's `[]`-suffixed parameters (`rule_name[]`, `rule_group[]`, `file[]`): vmalert reads only those and ignores plain `rule_name=`, an upstream quirk pinned by tests. `limit_alerts` is a Grafana parameter that vmalert ignores and is therefore omitted.

## Why the gate never queries state history

Querying Grafana's alert state history after the fact fails closed *in the wrong direction* — it returns "pass" when the truth is unknown:
Expand Down
8 changes: 7 additions & 1 deletion grafana-alertcheck/docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,12 @@ HTTP ──> Source ──> []StateRule ──> reduce ──> []Poll ──> pr
JSONL log ──> ReadLog ──┘
```

The `Source` interface is the only HTTP boundary and covers both rule kinds: `GrafanaDefinitions` reads the ruler, `DiscoverRuleSources` + `DatasourceDefinitions` read per-datasource Prometheus rules, and `RuleState` takes a `RuleRef` describing exactly how to find one rule again.

### Key vs uid

A rule's map key is its **key**, not its uid. For a Grafana-managed rule the key *is* the uid; a datasource-managed rule has none, so its key is a JSON `(datasource, group, name, file)` tuple — file included because a Prometheus group name is only unique within a file. Every internal map is indexed by key. The log records `key` and, for a Grafana rule, `uid`; the reader uses `rule_key` when present and falls back to `rule_uid`, so old v1 logs stay readable. Grafana behavior is unchanged because key == uid there.

- `proveCoverage` (the nine coverage checks) and `decide` (the instance timelines and outcomes) are pure; tests drive them with `[]Poll` literals and a fake `Clock`, with no sleeping or fixture server.
- `Check`/`Watch` are I/O shells: HTTP, signals, the pidfile, file reads, the countdown print. The only test doubles needed are the `Source` and `Clock` interfaces.
- `Policy` is the narrowed view of `Config` that reaches the pure layer — classification knobs and the window, no URL and no token. The token must never cross that line, which is the cheapest guarantee it never lands in an error string or a result.
Expand Down Expand Up @@ -76,4 +82,4 @@ On a clean stop (SIGTERM/SIGINT/`--until`) the child finishes the in-flight writ
`watch` records raw evidence, so nothing trusts a state that could become unreachable. Two consequences a maintainer must preserve:

- The **header is authoritative for recording facts** (the cadence actually used, the URL, the alert set); the ruler API is authoritative for **rule facts** (`for`, `intervalSeconds`, kind). `check` always re-resolves definitions fresh and never reconstructs them from the header — the header duplicates `for`/`interval` only so the uploaded artifact is self-describing.
- The **cadence authority** is the header's `poll_every_seconds`, not the definitions. Re-deriving it would compare gaps recorded at an override cadence against default-cadence thresholds — fail-open in the faster-override direction.
- The **cadence authority** is the header's `poll_every_seconds`, not the definitions. Re-deriving it would compare recorded gaps against thresholds derived from a different cadence — fail-open if the two ever diverge.
11 changes: 11 additions & 0 deletions grafana-alertcheck/docs/how-alerts-are-evaluated.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,17 @@ When an instance leaves the bad set, the gate looks it up **in the same response

A vanished instance that was bad stays `still_failing`. A metric that stops being emitted is not evidence of health — this is deliberate and can surprise users whose fix is to remove a metric rather than drive it to a good value.

### Datasource-managed rules

Datasource-managed (Prometheus-flavored) rules are read through Grafana's per-datasource Prometheus API, not the ruler. Their response carries only **active** instances, so an instance leaving the active set is treated as a real recovery (`cleared`). The weaker guarantee is documented deliberately: a vanished series is indistinguishable from a resolution, so a fix that stops emitting a metric passes for a datasource-managed rule where it would fail for a Grafana-managed one.

Two coverage checks differ for these rules:

- **Pause is not observable** — the datasource API has no `isPaused` signal, so check 7 is skipped with an explicit note. A pause is not treated as a pass; it simply cannot be seen.
- **Health** — the datasource vocabulary reports `err`, which is normalized to `error`, so a sustained failing evaluation still triggers check 4. There are no `totals`, reasons or normal instances, so checks 5 and 9 never fire.

There is also no `recovering` state: the datasource API keeps an alert `firing` through its *keep firing for* and then drops it, so the recovery observation below does not apply — an instance that clears simply leaves the active set, which is the `cleared` recovery described above.

## Coverage proof

Before classifying, `check` must **prove** continuous coverage of `[from, to]` for each alert. Nine checks run; any failure makes the rule `not_verified`:
Expand Down
2 changes: 1 addition & 1 deletion grafana-alertcheck/docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ export GRAFANA_TOKEN=…

Requires Grafana >= 13.0.0 and < 14.0.0. Outside that range the gate exits `2`.

Grafana API token needs to have `fixed:alerting:reader` permissions. Ask the o11y team for your token.
The Grafana API token needs `fixed:alerting:reader`, plus `datasources:read` and datasource query permission so datasource-managed rules can be discovered and read. Ask the o11y team for your token.

## Quickstart — recorder mode

Expand Down
Loading
Loading