Skip to content

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Checkpoint

A quality gate that sits between "draft" and "publish" in an automated content pipeline. It runs 12 checks on every draft. If a blocking check fails, the draft does not publish, it goes back for a rewrite with specific instructions, and gets re-gated.

Still fully automated. No human in the loop by default.

research  →  draft  →  ▓ CHECKPOINT ▓  →  publish
                            │
                            └─ fail → rewrite → re-gate   (max 2 passes, then park)

▶ Open the live demo (download it and open in a browser, or see it live once you turn on GitHub Pages) · Install · The 12 checks · Profiles · FAQ


Why this exists

Here is a site that published 1,635 blog posts in eleven months.

Publishing volume climbing while clicks flatline, then collapse

The bars are posts published per week. The line is clicks. Volume went up 21x. Clicks stopped responding in February and kept not responding for four months while the pipeline kept firing. Then, on one day in June, the site lost everything.

Of the 1,635 posts, zero were ever indexed.

Google Search Console page indexing report showing 1,635 pages not indexed

"Crawled - currently not indexed" means Google fetched the page, read it, and decided it wasn't worth putting in the index. Not a technical error. A judgment.

What Google actually penalises

Not AI. In Google's own words:

Producing content at scale is abusive if done for the purpose of manipulating search rankings, whether automation or humans are involved.

And:

Using AI doesn't give content any special gains. It's just content.

The quality rater guidelines are more specific. The lowest rating applies to content produced with little to no effort, little to no originality, and little to no added value, and the same section states plainly that "the use of Generative AI tools alone does not determine the level of effort or Page Quality rating."

So the useful question is not "was a model involved." It's: does this take effort, is any of it original, does it add value. Those three are measurable. AI detection is not.

You will probably never be told

In a documented 2026 experiment, two sites of roughly a thousand AI-generated posts each lost effectively all search visibility, one collapsing from 1,629 daily impressions to 15 overnight, with zero manual actions shown in Search Console. No email. No warning banner.

Enforcement at this scale is silent and algorithmic. Which means the gate has to come before publish, because there is no signal to react to afterwards.


Install (3 minutes)

Requirements: Python 3.8 or newer. That's it. No pip install, no API key, no network calls. A gate that needs a working install is a gate people skip.

Download this repo (green Code button → Download ZIP) and unzip it, or clone it:

git clone https://github.com/YOUR-GITHUB-NAME/checkpoint-quality-gate.git

Claude Code / Claude desktop

cp -r skill ~/.claude/skills/checkpoint          # available everywhere
# or, per project:
cp -r skill /path/to/your-project/.claude/skills/checkpoint

Then just say: "run the checkpoint on this draft."

Codex, Cursor, or any CLI agent

No skill loader needed. Point the agent at the doctrine file:

Read ./skill/SKILL.md and follow it. Gate ./drafts/my-post.md
with the b2b-saas profile before doing anything else.

Standalone, no agent at all

python3 skill/scripts/checkpoint.py drafts/my-post.md --profile b2b-saas

Verify it works

python3 skill/scripts/checkpoint.py skill/examples/failing-draft.md --profile b2b-saas ; echo "exit=$?"
# expect exit=2  (BLOCK)

python3 skill/scripts/checkpoint.py skill/examples/passing-draft.md --profile b2b-saas ; echo "exit=$?"
# expect exit=0  (PASS)

If the failing draft returns 0, the profile didn't load. The script will tell you so.


Running it

# basic
python3 skill/scripts/checkpoint.py draft.md --profile b2b-saas

# with the corpus check (the one that matters, see below)
python3 skill/scripts/checkpoint.py draft.md --profile local-services --corpus ./published

# machine-readable, for wiring into a pipeline
python3 skill/scripts/checkpoint.py draft.md --profile healthcare --json > gate.json

# write the report to disk
python3 skill/scripts/checkpoint.py draft.md --profile legal --report reports/draft.gate.md
Exit code Meaning
0 PASS: publishable
1 REWRITE: soft failures; rewrite and re-gate
2 BLOCK: a blocking check failed; must not publish
3 Error: bad path, unreadable profile, empty draft

Draft format

Markdown with optional frontmatter. The frontmatter is where you attach evidence, which is how waivers work (see Waivers).

---
title: What planned preventive maintenance actually costs
primary_keyword: planned preventive maintenance cost
author: Sam Okonjo
author_credentials: Operations lead, 11 years in commercial FM
last_reviewed: 2026-07-14
methodology_ref: internal-jobs-db-2023-2026
sample_size: 412
---

# What planned preventive maintenance actually costs

Planned preventive maintenance runs between £18 and £34 per asset per year...

The one rule

The gate must be a separate agent from the writer. Always.

A model grading its own draft passes it. Not through dishonesty. It is evaluating against the same understanding of the brief that produced the draft, so the draft looks correct by construction.

This is the single most common reason quality gates fail in practice, and it's invisible. You get a report full of green ticks and a corpus that quietly degrades anyway.

Three topologies, in order of preference:

Setup Use when
A Gate runs as a spawned subagent, synchronously, blocking Your tool supports subagents. Default.
B Gate runs in a fresh session; you paste the draft in and paste the verdict back No subagent support. Slower, equally valid.
C Draft parks as pending-gate and never reaches the CMS No second context available at all.

A parked draft is fine. A self-graded draft is worse than no gate, because it manufactures confidence.

When you spawn the judgment agent, give it the draft and the rubric and nothing else. If you pass it your reasoning about the piece, you've re-created self-grading with extra steps.


The 12 checks

Deterministic (Python, sub-second, no model call)

# Check Fails when
1 Depth Below the word floor, or H2s promising more than they deliver
2 Filler & effort Padding, stock phrasing and hedges above rate limits
3 Repetition & template collapse Internal repetition, or shingle overlap with what you already published
4 Substantiation A persuasive number or absolute claim with no basis in its paragraph
5 Compliance tripwires The industry profile's rule pack fires
6 Structure & accessibility WCAG 2.2 A/AA items checkable from source
7 Metadata & accountability No named author with credentials, no reviewer, no editorial-responsibility record

Judgment (graded by an agent that did not write the draft)

# Check The question
8 Question fidelity Does it answer the target query, in the first 150 words?
9 Original contribution Is there anything here that isn't already on page one?
10 Brand voice Would a regular reader believe the same team wrote this?
11 Experience signal Is there evidence anyone has actually done this?
12 Citability Can a passage be lifted out and still make sense and still be correct?

The script runs 1–7 and emits a JSON rubric packet for 8–12 at the end of its report. Hand that packet, the draft, and your brand voice file to the second agent.

Two checks worth understanding properly

Check 3 is the one nobody else runs. It compares every five-word sequence in the draft against every post in --corpus. A single post rarely looks like scaled content abuse. Forty posts sharing a skeleton with the nouns swapped is the textbook case, and it's invisible when you review drafts one at a time, which is how everyone reviews drafts.

Point it at your published folder. On a location-page corpus this is uncomfortable reading, and the discomfort is the finding.

Check 9 is the one that matters most and the one automation can't reach. List every claim a reader couldn't get from the top three competing pages. Each must be proprietary data, a named worked example with real numbers, a counter-consensus position argued with evidence, a synthesis nobody else has made, or a first-hand account of doing the thing.

"Explained more clearly" is not an original contribution. Neither is "more comprehensive."

Full definitions and tuning notes: skill/references/checks.md


Profiles: how the bar moves by industry

The checks stay constant. Thresholds and rule packs are data. Swap the profile, never the script.

Profile Built around
default FTC disclosure, review authenticity, AI overclaims, stale certifications
healthcare FDA disease and symptom claims, HIPAA in marketing, substantiation tier
finance FINRA 2210, SEC Marketing Rule, FCA standalone compliance
legal ABA Model Rules 7.1–7.3, jurisdiction-routed
b2b-saas Unbacked ROI, category-exclusivity claims, pricing and logo drift
ecommerce Affiliate disclosure, fake scarcity, review provenance
local-services Doorway-pattern location pages, review gating, NAP drift

Every compliance rule cites its primary source in skill/references/industry-profiles.md. A verdict a writer can't look up is a verdict they'll overrule.

Writing your own

Extend the closest profile and state only your deltas. Thresholds merge, rule packs concatenate.

{
  "name": "acme-health",
  "extends": "healthcare",
  "required_metadata": ["title", "author", "author_credentials",
                        "primary_keyword", "reviewed_by", "acme_ticket"],
  "thresholds": { "target_words": 2000, "min_original_claims": 3 },
  "compliance_rules": [
    {
      "id": "acme-vendor-neutral-clinical",
      "pattern": "\\b(?:CompetitorA|CompetitorB)\\b",
      "severity": "WARN",
      "why": "Internal policy 2026-03: clinical content stays vendor-neutral.",
      "fix": "Remove the brand name or move the comparison to a marketing page."
    }
  ]
}

Adding a rule? Four things, all required: why cites the actual rule or decision, fix is imperative and specific, near scopes it so it isn't bare keyword matching, and you test it against ten posts you're happy with before shipping. If it fires on any of them, it isn't ready.


Waivers

Every compliance rule supports unless_meta. A waiver requires attaching evidence, not just overriding.

---
title: Why our platform beats Competitor X on ingest speed
evidence_ref: bench-2026-07-14
evidence_date: 2026-07-14
---

That suppresses the comparative-claim rule, because the claim is now defensible, which was the point. A waiver with a reason is fine. A silently ignored warning is not.

Never waive by deleting a check. A disabled gate is invisible six months later; a loosened one still reports.


Wiring it into a pipeline

The gate is a shell command with a meaningful exit code, so it drops into anything.

Git pre-commit hook

Catches the drafts that skip your pipeline entirely, which is where most bad content comes from.

#!/usr/bin/env bash
# .git/hooks/pre-commit   (chmod +x)
set -u
fail=0
for f in $(git diff --cached --name-only --diff-filter=ACM | grep '^content/.*\.md$'); do
  python3 skill/scripts/checkpoint.py "$f" --profile b2b-saas --corpus content/published
  [ $? -eq 2 ] && { echo "BLOCKED: $f"; fail=1; }
done
exit $fail

GitHub Actions

A ready-to-use workflow ships at .github/workflows/content-gate.yml. Exit 1 (REWRITE) passes with a report attached; exit 2 (BLOCK) fails the build. No install step: that is the point of the zero-dependency rule.

Monthly corpus sweep, run this first

Usually the highest-value thing you'll do. Gate everything already published and sort by overlap.

for f in content/published/*.md; do
  python3 skill/scripts/checkpoint.py "$f" --profile local-services \
    --corpus content/published --json 2>/dev/null |
  python3 -c "
import json,sys
d=json.load(sys.stdin); g=[x for x in d['gates'] if x['gate']==3][0]
print(f\"{d['draft']},{g['detail']['corpus_similarity_pct']},{g['detail']['nearest']}\")"
done | sort -t, -k2 -rn | head -20

More patterns, including CMS webhooks: skill/references/deployment.md


Calibrate before you let it block anything

A gate that blocks 60% of drafts gets disabled in a week. A gate that blocks 8% gets trusted for years. The difference is thirty minutes of calibration.

  1. Take ten posts you're proud of and three you regret.
  2. Run the gate over all thirteen.
  3. Read what fired on the good ten. Each one is either a threshold that's too tight or a real flaw you hadn't noticed. Be honest about which.
  4. Read what did not fire on the bad three. This half is more valuable and everyone skips it. If your regrettable posts sail through, the gate is measuring the wrong thing for your content.
  5. Set thresholds so the good ten pass and the bad three fail. That's the whole calibration, your corpus, not an ideal, not a competitor's number.

Full guide, threshold by threshold: skill/references/calibration.md


What this deliberately does not do

  • It does not detect AI. Nothing does, reliably. The detectors on the market have false-positive rates that will fail your best human writers.
  • It does not throttle publishing volume. Volume isn't the problem. Publish daily if every post clears the bar.
  • It does not edit your draft. A gate that edits has become the writer, and now nothing is grading it.
  • It does not block competitor comparisons. The FTC actively encourages truthful comparative advertising and has criticised codes that hold comparative claims to a higher substantiation bar than unilateral ones. The real exposure is a Lanham Act suit from the named competitor, and the control is dated evidence, not suppression.
  • It does not call readability a WCAG AA requirement. Reading level is SC 3.1.5, Level AAA. It's reported as advisory and labelled that way.

FAQ

Does this work with ChatGPT / Codex / Cursor, or only Claude? Any of them. Checks 1–7 are a plain Python script with no model in the loop. Checks 8–12 need some agent, and any capable model will do, the rubric is in the report.

How long does a gate run take? Checks 1–7: under a second, even with a large --corpus. Checks 8–12: one agent call.

What if my drafts aren't markdown? Convert to markdown with frontmatter before gating. Most CMSs can export it, and the payload you send to the CMS afterwards is your business.

Won't 12 checks slow my pipeline down? The point isn't speed, it's not shipping things that cost you the rankings you built the pipeline to get. That said, the deterministic half adds under a second, and the rewrite loop is capped at two passes.

Everything is failing. What did I do wrong? You copied thresholds from a stricter profile. Calibrate against ten posts you're happy with. Also check the judgment agent isn't manufacturing faults to look rigorous, unevidenced failures are void, and the rubric says so explicitly.

Everything is passing. What did I do wrong? Almost always: the writer is grading itself. See The one rule. Second most likely: --corpus isn't actually pointed anywhere.

Can I use this commercially / with clients? Yes. MIT licensed. Fork it, rebrand it, sell it as part of your service.


Repo layout

README.md                          this guide
docs/
  index.html                       GitHub Pages landing page
  demo.html                        the interactive demo
skill/
  SKILL.md                         doctrine (the agent reads this)
  scripts/checkpoint.py            checks 1–7, zero dependencies
  profiles/*.json                  industry rule packs
  references/checks.md             all 12 checks in full
  references/industry-profiles.md  the primary source behind every rule
  references/calibration.md        how to tune without breaking trust
  references/deployment.md         Claude Code, Codex, git hooks, CI, CMS
  templates/brand-voice.template.md
  examples/                        a failing draft, a passing draft
.github/workflows/content-gate.yml   optional: gates content on every pull request

Publishing your own copy

One repo. Everything lives here. The skill, the guide and the demo ship together, because the guide's whole job is to get someone to the skill, and a reader who has to go find a second repo mostly does not.

No terminal needed:

  1. On GitHub, click + (top right) → New repository. Name it checkpoint-quality-gate, set it Public, and do not tick "Add a README file". Click Create repository.
  2. On the empty repo page, click uploading an existing file.
  3. Drag in everything from this folder: README.md, LICENSE, and the docs and skill folders. Wait for all of them to finish, then click Commit changes.
  4. Go to Settings → Pages. Under Source pick Deploy from a branch, set branch to main and folder to /docs, and click Save.
  5. Wait about a minute, then open https://YOUR-GITHUB-NAME.github.io/checkpoint-quality-gate/

Nothing needs editing afterwards. The guide works out its own links.

Two URLs come out of that one repo:

URL What it is Who you send it to
github.com/you/checkpoint-quality-gate The repo. README renders as the landing page. Developers, anyone who will clone it
you.github.io/checkpoint-quality-gate/ The guide, styled, with a live demo Everyone else, DMs, link in bio

Both are the same commit. Push once, both update.


Legal note

skill/references/industry-profiles.md is built from primary sources and dated August 2026. Regulations move: FINRA proposed modernising Rule 2210 in July 2026; the FTC's Click-to-Cancel rule was vacated in July 2025 and rulemaking restarted.

This is not legal advice. In a regulated vertical, have counsel review your profile once before it goes live. That review is cheap. The alternative is not.

Sources

Google spam policies · Quality Rater Guidelines · Google on AI-generated content · FTC Endorsement Guides, 16 CFR 255 · FTC Reviews Rule, 16 CFR 465 · FDA 21 CFR 101.93 · FINRA Rule 2210 · SEC Rule 206(4)-1 · ABA Model Rules 7.1–7.3 · WCAG 2.2 · EU AI Act Article 50

License

MIT. See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages