Give your AI a task. Get it back with evidence for every claim.
Do the Thing is an installable agent skill that connects source review, interviews, Big and Mini Painless Specs, plan critique, execution, and recovery. The name comes from the Korean phrase for “do some work.”
Available now: the skill package with its interview contract, Big/Mini/checkpoint templates, a completion-record validator whose suite is mutation-tested in CI, fictional examples, and evals/ — a re-runnable suite that scores the skill against the same work done without it. A tracker is optional — Linear has a contract here, another tracker follows the same rules, and without one a file in the repository is the work record. The measured benefit was produced with no tracker connected. This is not a hosted app, OAuth integration, background worker, or autonomous webhook service. Five runs so far, two of them on live Linear issues — one closed to Done with the whole journey verified and resume exercised, one still in progress. Native comment threads and crash-simulated resume are untested. See release status.
npx skills add beamonic/do-the-thingThat is the whole install for any agent the skills CLI supports.
Version 0.20 makes the tracker optional — the workflow runs unchanged
without one — and it is the first change whose pass criteria were committed
before the runs that judged it. evals/ holds that loop: a re-runnable suite
that scores the skill against the same work done without it, so a rule change
can be shown to help or hurt instead of asserted. Seven rounds have run, and
each of the first six found a fault in the measurement before it found one in
the skill.
See release notes
and what the numbers do and do not support.
To place the folder by hand instead — for Codex:
git clone https://github.com/beamonic/do-the-thing.git
mkdir -p ~/.codex/skills
cp -R do-the-thing/skills/do-the-thing ~/.codex/skills/If that destination already exists, compare it before replacing your installed copy. Other agents that support SKILL.md can use the same self-contained folder in their own skills directory. The frontmatter deliberately stays inside the six keys the Agent Skills spec allows, so a Claude Code extension such as when_to_use is not used and tests/test_skill_frontmatter.py fails if one appears.
The description leads with the trigger phrases rather than the summary. A host carrying many skills shortens every description to roughly its first sixty characters — measured at 60–64 on a Codex host with 154 skills installed — so whatever has to survive that cut goes first, and what has to survive is the matching. English triggers come first and the Korean ones are kept; the skill is used in both. The value stays a plain YAML scalar: a value that begins with a quote is a quoted scalar, and a strict loader rejects whatever follows its closing quote. tests/test_skill_frontmatter.py fails on that shape, on a colon-space, and on fewer than four distinct triggers.
Connect Linear through your agent's supported connector, MCP, or API client. The skill discovers the available operations; it does not bundle credentials or assume a specific tool name. Python 3.9+ is needed only for the optional record validator; it is tested on 3.9 and 3.13.
Use $do-the-thing on Linear issue TEAM-123.
Read its sources, interview me about unresolved decisions, and create a Big
Painless Spec. Split it into task-level Mini Painless Specs, then execute
work within the permissions I have given you. Keep evidence and checkpoints
in the same work context so we can resume later.
TEAM-123 is a placeholder. Start with a non-sensitive trial issue. If Linear is not connected, the agent can prepare drafts locally but must not report them as synchronized.
You only need to ask: 두더띵으로 이 일 해줘 (Use do-the-thing for this task).
The agent selects the relevant internal method: domain modeling for conflicting
terms, test proof for uncertain coverage, and spec/standards review for meaningful
code changes. These three contracts ship inside the skill folder; no separate
commands or skill installations are required. It reads only the methods needed
at the current step, reuses accepted decisions, and skips heavy work for simple
requests. A request to review or diagnose remains read-only. This is instruction
routing, not a new API or a background automation service.
Use do-the-thing as the single entrypoint. Its Ponytail-inspired principle
checks necessity and reuse throughout the grill, spec, research, and execution:
reduce unnecessary work, not outcome responsibility. Prefer an adequate existing
solution before adding machinery; consider maintenance and burden shifted to
others, not only code size. Explicit requirements, safety, and whole-journey
verification remain intact. Ponytail is adapted here, not an extra command,
dependency, intensity setting, or persistent mode.
Version 0.16 accepts an existing grill, specification, and plan at the first unfinished step. It preserves accepted decisions, IDs, revisions, and execution authorization. It does not require another interview or a rewritten spec solely to fit its templates. Damaged checkpoints are preserved during recovery; current artifacts and every Big criterion still need verification before completion.
Mutation tests now run in a temporary copy and reject zero-test runs. See release status and terminal evaluation.
Does this issue need a spec at all? → no: do the work, record it, stop
yes:
Sources → Interview → Big Painless Spec → Plan and critique → Split tasks
For each task:
Mini Painless Spec → Focused follow-up interview → Plan and critique
→ Execute → Verify Mini → Checkpoint → Next task
Then:
Verify the whole Big user journey → Deliver and read back in Linear
| Spec | Responsibility |
|---|---|
| Big Painless Spec | Overall user outcome, scope, shared rules, exceptions, and acceptance criteria |
| Mini Painless Spec | One verifiable task, its behavior and evidence, linked to the Big revision and criteria |
The first step is deciding whether to run the rest. An issue that is already clear, bounded, low-risk, and names its own change gets done directly, with the result recorded in the issue — no Big, no Minis. The workflow's likeliest failure is ceremony on a task that never needed one, not a missing spec.
An interview is an explicit step, and its contract ships with the skill. Questions are asked one at a time from the frontier — the decisions whose prerequisites are already settled — each with the agent's recommended answer, and the agent waits through a permitted native blocking question tool. Codex requires a mode that permits request_user_input; asynchronous questions plus sleep are not a substitute for a pending input request. An answer is transcribed under a stable question ID before anything is acted on; a question nobody answered stays open, and nothing is built on it. Facts the agent can look up are not interview questions. The interview may end with questions still open when they affect only individual tasks. The frontier discipline is adapted from Matt Pocock's grilling skill, which asks a whole frontier per round; this skill asks one question per round because a batched list with recommendations was getting answered as a whole or not at all. Nothing needs to be installed alongside this one. A Mini inherits shared rules instead of asking the whole interview again. If a task uncovers a change to the overall promise, revise the Big and review affected work.
If the repository already carries GitHub's Spec Kit, the agent calls its speckit.* commands at each of those steps — clarify feeds the interview, specify, plan, and tasks write the artifacts, implement runs one Mini at a time, analyze and converge feed verification — as the Spec Kit contract maps them. Nothing is installed for you, and the workflow is the same without it.
All Minis passing does not mean the Big passes. Verify the complete user journey before closing the work. A completed agent session is not a completed Linear issue.
When oreilly-context is available, the agent proactively uses it for knowledge
gaps before the grill, design trade-offs during the spec, and unfamiliar
implementation or repeated failures during execution. It skips mechanical work,
reuses relevant searches, and distinguishes candidate metadata, sections actually
read, and changes verified locally. The integration is optional; no subscription,
account, or research skill is bundled.
python3 skills/do-the-thing/scripts/records.py examples/completion.json
python3 -m unittest discover -s tests -vGitHub Actions runs both on every push and pull request: the suite on Python 3.9 and 3.13, and a mutation gate.
python3 scripts/check-mutations.pyThe gate deletes each of the validator's guards in turn and requires the suite to fail. A guard no test covers is a guard that can be removed without anyone noticing, which is what happened here before 0.6 — every test asserted only that some error appeared, so an unrelated guard firing on the same fixture covered for the missing one. tests/mutations.json lists the guards; renaming one without updating that list fails the gate rather than passing quietly.
The example is fictional. Each Mini criterion names the Big criterion it serves, and a Big criterion counts as covered only when a Mini criterion aimed at it actually passes. The validator checks coverage, declared revisions, and evidence fields; it cannot prove that the evidence is true. It never contacts Linear or changes an issue's state. It does warn when a record still carries the example's fictional flag or placeholder evidence, but a warning does not fail the record: read the warnings before trusting a pass.
- Skill entrypoint
- Interview contract
- Spec Kit contract
- Product spec and implementation boundaries
- Method and attribution
- Fictional demonstration
- Release status
- References
- Publication draft
Run the tests above. For behavioral changes, include a realistic input, the expected decision, and evidence of the result. Do not include real customer issues, private documents, or credentials in examples. Report whether a result came from a fixture, a local execution, or a live integration.
MIT. This project is independently authored. It is not affiliated with Linear, OpenAI, GitHub, Gajae Code, or Matt Pocock. No third-party skill implementation or runtime is bundled; host integrations are optional. The interview contract adapts the frontier discipline of the grilling skill (MIT, Copyright (c) 2026 Matt Pocock), rewritten here; see method and attribution.