A compliance engine for Azure DevOps. You describe the organisation you want in readable YAML — products, teams, area paths, repos, pipeline folders, permissions, iteration cadence — and the engine continuously asserts that the live Azure DevOps project matches it.
Config is truth. Live state must match config. Any difference is a violation.
This is deliberately not a provisioning helper that creates what is missing and moves on:
| What this is | What this is not |
|---|---|
| Compliance engine — drift is a finding | One-shot provisioner — "good enough" is not |
| Idempotent reconciler — every run converges | Set-and-forget — trust that it worked once |
| Corrective — a misconfigured team gets patched | Additive — only creating what does not exist |
| Auditable — every deviation is reported | Silent — unexpected state is ignored |
After a successful apply, an audit returns zero findings. That invariant
is the whole point.
Large Azure DevOps projects rot. Teams get created by hand with the wrong area configuration, someone is added to a security group and never removed, a repo ends up writable by the wrong team, and a year later nobody can say whether the project reflects the org chart or a series of accidents.
Point-and-click governance cannot be reviewed, diffed, or rolled back. This project makes the intended structure a reviewable artefact in version control, and makes every divergence from it visible and correctable.
You do not clone this repo to use it. This repo is the engine. You create
your own private workspace repo, declare this engine as a capability, and a
loader materialises the module into a generated .system/ folder. Your desired
state stays in your repo; the engine stays in this one. Nothing is duplicated and
nothing of yours is ever committed here.
Prerequisites
- PowerShell 7.4+
- Azure CLI with the
azure-devopsextension (az extension add --name azure-devops) - GitHub CLI for the repo-creation step
gh repo create my-org/my-governance --private --clone && cd my-governancecurl -o init.ps1 https://raw.githubusercontent.com/nkdAgility/azure-devops-automation-tools/main/system/NKDAgility.AzureDevOps.AutomationTools/Templates/customer-repo/init.ps1Create capabilities.json — this file is yours and is never overwritten:
{
"capabilities": [
{
"name": "automation",
"module": "NKDAgility.AzureDevOps.AutomationTools",
"repo": "https://github.com/nkdAgility/azure-devops-automation-tools.git"
},
{
"name": "governance",
"module": "NKDAgility.AzureDevOps.Governance",
"repo": "https://github.com/nkdAgility/azure-devops-governance-as-code.git"
}
]
}pwsh -NoProfile -Command ". .\init.ps1"That clones each engine, copies its module into .system/, scaffolds the rest of
the workspace (.gitignore, secrets/, governance/, agent guidance), and
exports your secrets as environment variables.
Create governance/programs/<name>/ (see Defining a
program), then — everything except apply is read-only:
az loginpwsh -NoProfile -Command ". .\init.ps1; Invoke-Governance build myprogram"
pwsh -NoProfile -Command ". .\init.ps1; Invoke-Governance plan myprogram"
pwsh -NoProfile -Command ". .\init.ps1; Invoke-Governance apply myprogram -WhatIf"
pwsh -NoProfile -Command ". .\init.ps1; Invoke-Governance apply myprogram"
pwsh -NoProfile -Command ". .\init.ps1; Invoke-Governance audit myprogram"Migrating teams in from elsewhere? Preflight gives each one its fix list before it moves.
Always run through a fresh shell so init.ps1 executes first. The engine runs
from .system/, which is only refreshed when init.ps1 runs — a shell that has
been open all day keeps serving yesterday's engine even after a git pull.
-NoSync skips the pull for offline work.
init.ps1 scaffolds secrets/secrets.example.json and gitignores everything else
in secrets/. Each entry names the environment variables to export, and your
manifest.yaml references one by name, never by value:
auth: entra # default: the signed-in az identity
accessToken: $Env:AZDEVOPS_MYORG_PAT # fallback for non-interactive runsYour workspace .gitignore keeps the generated and secret parts out of git:
/.system/ # materialised engine copies — generated, never committed
/secrets/* # PATs
!/secrets/secrets.example.json
/output/ # build artefacts and audit reports
workspace.local.jsonThere are exactly two setups, and the difference matters for whether a compliance run is reproducible.
| Consumption | Development | |
|---|---|---|
| Source | Published module, pinned version | A clone (yours, or a fork's) |
| Changes propagate | On an explicit version bump | Instantly, uncommitted edits included |
| Selected by | Default | enginePaths.<name> in workspace.local.json |
| Reproducible | Yes | No — came from one machine's working tree |
Consumption is the default and the only mode a scheduled audit should ever run in. Install it directly:
Install-Module NKDAgility.AzureDevOps.Governance -Scope CurrentUserDevelopment is for changing the engine, where edits have to take effect in a
real workspace immediately — no publish step in the loop. That is why a clone,
not a package: init.ps1 copies the working tree, uncommitted changes and all,
and warns you when it does. Point a workspace at your clone in
workspace.local.json, which is gitignored, so the override never reaches your
teammates or CI:
{ "enginePaths": { "governance": "C:\\src\\azure-devops-governance-as-code" } }Contributors work the same way against a fork — it is the development setup with a different remote, not a third mode. Clone the fork, point a workspace at it, change engine and configuration together, then open a PR.
Working on the engine on its own, without a workspace, use the CLI directly:
pwsh ./build.ps1 audit -Program myprogram -ProgramsRoot ../my-governance/governance/programsProvenance. A run in development mode is not evidence of anything reproducible — it came from a working tree that may exist on exactly one machine.
.system/.source.jsonrecords the version, commit, and whether the source was dirty. Check it before treating an audit result as authoritative.
Consumption mode defaults to the production ring, which has nothing on it
until the first stable release. init.ps1 does not fail the workspace over that.
It never tears down a working engine without a replacement in hand:
.system/already has the engine — it is kept exactly as it is, and its.source.jsonis left untouched. The workspace keeps running what it was already running, and its provenance still names whatever produced it.- Otherwise, a version is installed — that is staged instead.
- Otherwise it fails, naming the module, the ring, what is actually published, and both ways out.
Step 1 deliberately beats step 2: moving a workspace onto an unrelated version that happens to be in the module path, because a release hasn't happened yet, would change what it runs without anyone asking — worse than being stale. Every fallback is warned about, because it means the run is not on the ring it asked for.
| Command | What it does | Writes to Azure DevOps |
|---|---|---|
build |
Compile programs/<name>/*.yaml → out/<name>/resolved.yaml. Validates schema, unique codes, owner refs, Azure DevOps path limits. No live calls. |
No |
validate |
Same checks as build, writes nothing. |
No |
plan |
Diff the resolved desired state against live Azure DevOps. Lists every change apply would make. |
No |
apply |
Reconcile live Azure DevOps to the desired state. Creates missing resources and corrects misconfigured ones. Supports -WhatIf and -Prune. |
Yes |
audit |
Read-only compliance report. Reports every resource that deviates — missing, extra, or wrongly configured. | No |
preflight |
Per incoming team: if this team moved in today, what would fail? Reads its pre-migration location from sources.yaml and evaluates it with the same rules audit uses. |
No |
preflight-report |
Render each gathered team's markdown fix report from the files preflight wrote. Offline and deterministic — no organisation is contacted. |
No |
doctor |
Probe the current identity's permissions against the live org with side-effect-free calls. Exits non-zero listing what is missing. | No |
# Non-compliant live state # what apply does
Team "Foundation" exists ✓ no create needed
but areaConfig = [\Odyssey] ✗ wrong — patches it
Group "PTL-FND-Contributors" exists ✓ no create needed
but is missing alice@corp.com ✗ wrong — adds her
but also contains bob@corp.com ✗ wrong — removes him
By default, resources that exist in Azure DevOps but are absent from config are
reported as audit exceptions, not deleted. apply -Prune (or settings.prune: true in the manifest) additionally deletes orphan teams, area paths, repos,
and extra group members. Placeholders marked scope: future, the project default
team and repo, and iteration paths are never pruned.
audit answers "does this organisation match its config?". preflight
answers a different question, about a team that has not arrived yet: if this
team's content landed in the governed project today, which checks would fail?
That is the question that matters when a programme is consolidating. Teams have to reshape their area paths, tags and membership in their current organisation, before they migrate — so each one needs its own concrete list, not a general standard to interpret.
Preflight reads a team's pre-migration location, declared per node in
sources.yaml, projects that state into target coordinates, and evaluates it
with the same evaluators audit uses. The two can therefore never disagree
about what compliant means.
# programs/<name>/sources.yaml
sources:
PTL-FND:
org: legacy-org
project: LegacyPortal
areaPath: LegacyPortal\Foundation # project-rooted, no leading backslash
teams: ['Foundation Crew'] # optional: who works there today
repos:
include: ['Foundation*'] # optional: which repos are theirsIt is read-only against both organisations, and exits non-zero on findings,
the same CI contract as audit.
A long-lived area is mostly archive. Validating tags and placement across work
items nobody will migrate hands a team a fix list it should not be asked to
act on, so scope: narrows the population to the migration query — the
programme's own decision about what comes across:
scope:
label: "2026.1 and onwards"
query: >-
SELECT [System.Id] FROM WorkItems WHERE [System.TeamProject] = 'Proj'
AND [System.IterationPath] NOT UNDER 'Proj\ARCHIVE'It is a complete flat WIQL query. The query owns project and source area
selection, so it may select work from multiple areas. The engine adds ID paging
and replaces any authored ordering with System.Id ASC while gathering. A node
may override the programme default with its own scope:. Selected items outside
the configured source area root are reported as unmapped until their target
placement is defined.
This is deliberately the same text your migration toolchain needs for its own work-item query, so one authored decision drives both. A saved Azure DevOps query is not accepted as the input — it is mutable state outside version control, and a report whose population depends on one is not reproducible evidence.
The query is recorded in the gathered data and named in the report header, and changing it re-gathers rather than reusing a data file that describes a different population. Run without it and the report says so — "every work item under the area, archive included".
The run splits in two, and the split is the point:
| Step | Reads | Writes | Network |
|---|---|---|---|
| Gather | both organisations | -data.json — facts only, no verdicts |
Yes |
| Analyse | that data file | -findings.txt and -findings.json |
No |
| Render | data + findings | -report.md, the team's fix report |
No |
-Offline re-analyses the last gathered data, so tightening a rule costs
nothing and needs no credentials. -SkipFresh gathers only the nodes that have
no data yet — the resume path when a Conditional Access sign-in expires part
way through a large programme.
Artefacts land one folder per node, named so they still identify themselves after being filed somewhere else:
<output>/preflight/
<program>-preflight-summary.md
<CODE>/
<program>-preflight-<CODE>-data.json
<program>-preflight-<CODE>-findings.txt
<program>-preflight-<CODE>-findings.json
<program>-preflight-<CODE>-observations.md
<program>-preflight-<CODE>-report.md
preflight-report renders the entire markdown document, and every count and
every table in it is copied from the two input files. The only prose anyone
else contributes is -observations.md, spliced into one bounded section. So
whatever writes the judgement — a person, or a model — cannot reach a table and
cannot get a number wrong.
Findings are objects rather than strings: class, a stable check id,
subject, counts, examples, plus whatever your labels: map attaches per
check. That is where your own team-facing standard plugs in — a rule number, an
owner lane, a task id. The engine knows nothing about your document; your
program maps the engine's checks onto it.
labels:
area.orphan: { rule: A2, task: 2, lane: PM }
tag.unsanctioned: { rule: B4, task: 6, lane: Eng }
reporting:
candidateTagMinUses: 20 # the bar every team is measured againstA tag that is not in the vocabulary still has to go somewhere, and which
somewhere decides who does the work. taxonomy.yaml records the decision:
tags:
sanctioned: [...] # names what the work IS — it stays
boardColumns: [...] # names where work has GOT TO — becomes a column
retire: [...] # nothing depends on it — delete
sanctionedPatterns: # a family minted fresh each season
- pattern: '^\d{4}S\d+(Committed|Stretch)$'
note: "moving to iteration paths as teams adopt the new project""Test passed" and "Kicked off" are board columns wearing a tag's clothes; sanctioning them entrenches the workaround, and a column shows the stage at a glance instead. The report lists board columns, retirements and undecided separately, so the question a team answers is "which of three" rather than "what do you want to do about 549 tags". A tag given two destinations fails the build.
There is deliberately no "tolerated" state. The audit knows only the
sanctioned vocabulary, so a tag nobody can stop being applied — one stamped by
tooling outside the team's control — must be sanctioned, or it is an exception
every day forever and a -Prune deletion candidate.
A legacy area can carry thousands of machine-generated tags — build ids, crash
session ids. Declare them as disallowedPatterns in taxonomy.yaml and each
family collapses to one finding, carrying its distinct-tag count, its
work-item count and examples, instead of one finding per tag. What is left is
the vocabulary a team actually has to decide about.
The engine scaffolds an operator workflow into your workspace under .claude/,
all engine-managed: an /audit-preflight command, the workflow it runs, two
subagent definitions, a skill holding the observation rules, and a PreToolUse
hook that refuses any shell command that would run apply. Cheap models do the
gather, render and publish; one capable model per team writes the observations;
a cheap checker verifies that every number in them exists in the data.
None of that is required. preflight and preflight-report are plain
PowerShell, so CI or a bare shell produces the same reports for every team,
lacking only the observations section.
A program is a folder of YAML. Only manifest.yaml and hierarchy.yaml are
required.
governance/programs/<name>/ # in YOUR workspace repo, not this one
manifest.yaml # program identity, org, project, auth reference
hierarchy.yaml # the authored product / structural / team tree
access.yaml # role definitions + group naming conventions
taxonomy.yaml # governed vocabularies (work item tags) — optional
systems.yaml # reusable team sub-elements — optional
cadence.yaml # iteration cadence — optional
sources.yaml # pre-migration locations, for preflight — optional
members/ # <code>.yaml — desired group membership
manifest.yaml — identity and connection. accessToken is always an
environment-variable reference, never a literal token:
program: Odyssey
org: my-org
accessToken: $Env:AZDEVOPS_PAT # only used when auth: pat
project:
name: Odyssey
process: Agile
visibility: private
sourceControl: githierarchy.yaml — position in the tree is the type. A node is a product
because it has a dpm; a node with child teams is structural; a leaf is a
delivery team.
products:
- name: Portal
dpm: 101
short: PTL
pipelineFolder: true
sections:
- name: Platform
items:
- name: Graphics Pipeline
short: GPI
systems: [bug-inbox]
- name: Foundation
short: FND
type: delivery # a working team even though Open API nests under it
iterations: none # a delivery team that does not plan in sprints
- name: Plugins
items:
- name: Plugin A
short: PLGA
sideload: PTL-FND # area path here, on Foundation's board. No security.
repos:
- plugin-a # -> PTL-PLGA-plugin-a (auto-prefixed with the node code)A node's global key is the product-qualified chain of short codes — Portal's
Graphics Pipeline is PTL-GPI. That key names its groups (PTL-GPI-Contributors)
and its membership file (members/PTL-GPI.yaml).
A section normally groups its items without creating a team of its own.
When the section itself is accountable, give it an explicit type and short:
sections:
- name: Engineering
type: delivery
short: ENG
pipelineFolder: true
repos:
- ToolsUnder Portal (short: PTL), this creates one Portal\Engineering area and
one PTL-ENG team, owning PTL-ENG-Tools. Its membership still comes from
members/PTL-ENG.yaml; its area authority and pipeline folder use the single
area. No repeated child is needed. type accepts delivery, structural
or portfolio, with the same planning defaults as other typed nodes.
Optional section children use items, not teams; their codes extend the
section code. Sections without type keep their existing grouping behaviour.
When flattening an existing nested team, retain its short code, repositories, membership file and team-ID mapping. The area and pipeline-folder paths change, so review the desired-state difference before any later live application. Building and validating this configuration are offline operations; they do not move or rename live resources.
members/<code>.yaml — desired membership, reconciled in both directions.
Every entry carries a reason, so an access grant is self-documenting:
contributor: []
reader: []
admin:
- upn: jordan.blake@example.com
reason: "Structural authority over the Foundation subtree."sideload— area-path visibility only. The path joins the listed teams' boards and grants no permission whatsoever. Structural authority always stays with a team's home area.adminrole — delegated structural authority over a subtree (its area paths, teams, and settings). It is not team membership and grants no code access.scope: future— a structural placeholder. Invisible toapplyandaudit. Remove the flag when the product enters active migration.systems— reusable sub-elements (e.g. abug-inbox) stamped onto a team as child areas of its home area. Uniform across every team that applies one, and carrying zero security.
Entra is the default. Runs authenticate as the signed-in az identity —
interactively via az login, or in CI via OIDC. Entra tokens carry the
identity's real permissions: no PAT scopes to configure, no secret to store, and
they work in organisations that forbid full-access PATs.
Mode resolution:
- manifest
auth: pat→ use theaccessTokenPAT (an explicit opt-out of Entra) - otherwise, an
azsession is present → entra - otherwise,
accessTokenresolves → pat, with a warning (CI fallback) - otherwise → error, telling you to
az login
Whichever mode you are in, run doctor first. It probes every permission family
against the live organisation using reads plus intentionally invalid writes that
Azure DevOps rejects after the permission check — an HTTP 400 proves access,
and nothing is ever created.
If you must use a PAT, the required scopes are listed in
CONTRIBUTING.md. Note that security ACL
writes need permissions no PAT scope short of full access grants; the engine
falls back to an Entra token from your az login session for those.
build.ps1 # CLI entry point, for working on the engine directly
programs/ # always empty here — program definitions live in the
# consuming workspace repo, never in this one
system/NKDAgility.AzureDevOps.Governance/
Private/Compile/ # build stage: Import -> Resolve -> Write/Test
Private/AzureDevOps/ # thin wrappers over the Azure DevOps REST API
Private/Compliance/ # the reconcile loop shared by plan/apply/audit
Public/ # exported cmdlets, one per command verb
Templates/ # what gets scaffolded into a consuming workspace
Agents/ # guidance for AI agents using this as a capability
.agents/ # guidance for AI agents and contributors working ON the engine
decisions/ # architecture decision records
tests/ # Pester tests
out/ # generated artefacts (gitignored)
The reasoning behind the non-obvious choices is recorded in
.agents/decisions/:
| ADR | Decision |
|---|---|
| ADR-001 | PowerShell over Go — right for internal tooling |
| ADR-002 | Targeted REST checks over bulk fetches — bulk listing causes OOM on large projects |
| ADR-003 | Apply is corrective, not just additive |
| ADR-004 | scope: future nodes are structural placeholders, filtered from apply and audit |
| ADR-005 | Old governance iterations are compliant by definition |
| ADR-006 | Sanctioned tags are made to exist via an anchor work item |
| ADR-007 | Preflight evaluates projected pre-migration state through the audit's own evaluators |
| ADR-008 | Preflight gathers a facts-only data document first, then analyses it offline |
| ADR-009 | The renderer owns the fix report; agents write one observations fragment |
| ADR-010 | The migration query scopes preflight; the 2026-09-17 amendment accepts a complete WIQL query with multiple source areas |
| ADR-011 | A tag that is not sanctioned still needs a destination: board column, retirement, or a decision |
Contributions are welcome. Start with CONTRIBUTING.md — it covers the two non-negotiable rules (run the tests; never fail silently), the checklist for adding a new governed resource type, and how to run the suite:
pwsh -NoProfile -Command "Invoke-Pester ./tests -Output Normal"Participation is governed by our Code of Conduct. To report a security issue, see SECURITY.md — please do not open a public issue.
GNU AGPL v3 © naked Agility Ltd (Martin Hinshelwood & Co.)
If you run a modified version of this engine as a network service, the AGPL requires you to offer that modified source to its users.