Status: v0 skeleton, not yet announced. This is an early, partial cut of a larger internal framework — see "What's here" and "What's deferred" below before assuming this is a complete, ready-to-run system.
An AI-native framework for running a company's day-to-day operations with a small hierarchy of agents, a governed department-loop pattern, a human-in-the-loop (HITL) approval layer, and a learning-from-outcomes feedback loop — built on the Cursor SDK as the reasoning backend.
This is the generic "how an AI runs a company's operations" mechanism, deliberately extracted from a specific company's real deployment and stripped of that company's actual business logic, credentials, and strategic data. It is a framework/template, not a product — you bring your own department loops, your own doctrine docs, and your own credentials.
| Path | What it is |
|---|---|
learning/ |
Outcome tracking (win/loss signal mining from logs), tactic memory (few-shot retrieval over outcome history), skill extraction (codify recurring successful agent tasks into reusable skill files), leadership track record (did leadership recommendations actually help?) |
hitl/ |
Human-in-the-loop packet queue helpers: queue_io.py (enqueue to markdown + JSONL, never execute), downstream outcome tracking, dual-control (two-person integrity) gate for maximally dangerous actions, HMAC-signed magic links for off-network phone approve/deny |
hands/agent_runtime.py |
A Cursor-SDK "hands" worker: drains a file-based task queue, runs bounded Agent.prompt() calls with a read-only web-fetch tool (SSRF-guarded), per-task and per-day token ceilings, and skill injection |
cursor_bridge.py |
A lightweight fire-and-forget Cursor SDK wrapper (the "shape 4a" call every other module here uses for a single reasoning/drafting call), with tiered model routing and a token ceiling |
goal_engine.py |
Translates a plain-English yearly plan into small, clamped (±50%) KPI-target deltas for your own department loops — never touches a hard-stop or approves/denies anything |
leadership_common.py |
Shared plumbing for a chief-agent hierarchy (CEO/COO + domain chiefs): doctrine-excerpt reading, Cursor-prompt wrapper, bounded sub-agent delegation onto the hands queue |
- The chief-agent tick pattern itself (
ceo_agent_tick.py,coo_agent_tick.py, domain-chief scripts) — the pattern is generic but each real implementation has doctrine-path and approver-name specifics that need per-file genericization, not a blind copy. - A worked, fully-generic example department loop (the
HARD_STOP/KPI/HITL_FORCE/ audit-log template every department loop should follow) — recommended as one clean example rather than porting many business-specific loops. - A tech-scouting module (bounded, keyless, real-API-only "verify before recommending" pattern) — reusable in shape, needs its example domain config genericized.
- An output-grading loop (grade artifacts against a rubric, delegate at most one bounded fix task) — the mechanism is generic; needs a genericized example rubric.
- A growth-trajectory scoreboard pattern (informational-only, explicitly never an optimization target) — needs real numbers stripped and the reference curve made a parameter.
- The full HITL packet-queue engine + chat-card rendering (larger files, deserve a careful pass rather than a rushed one).
- Generalized versions of the underlying design docs (leadership authority model, KPI-nudge guardrail, department-loop contract).
- Bounded, never a new lever. Every automated write (a KPI target nudge, a sub-agent delegation, a magic-link mint) is clamped, capped, and logged — nothing here grants an agent a new class of unsupervised authority.
- Honest gaps, never fabricated. A module that can't do real work yet
(no API key, no real plan file, insufficient log volume for a stats
signal) reports that state explicitly (
insufficient_data,blocked_on_cursor, etc.) instead of guessing or inventing a plausible- looking result. - Fail-closed. Every Cursor SDK call site checks availability first and
degrades gracefully (heartbeat-only,
None, an explicit blocked status) rather than crash-looping when a key isn't configured. - HITL stays a human's call. Dangerous actions get a real approval gate (a single-approver HITL queue, or two-person dual-control for the most dangerous action classes) — this framework enqueues and tracks decisions, it does not make them.
- Everything is logged. Every module appends to its own append-only JSONL audit log — this is a design choice worth keeping if you extend it.
pip install cursor-sdk httpx
# Cursor API key (either works):
export CURSOR_API_KEY=your-key-here
# or:
mkdir -p ~/.config/phantom-box && echo "CURSOR_API_KEY=your-key-here" > ~/.config/phantom-box/cursor.env
chmod 600 ~/.config/phantom-box/cursor.env
# Point every module at your own working directory (defaults to ~/phantom-box):
export PHANTOM_BOX_ROOT=/path/to/your/deploymentpython3 cursor_bridge.py --selfcheck
python3 hands/agent_runtime.py --selfcheck
python3 learning/outcome_tracker.py --dry-run
python3 goal_engine.py --dry-runMIT — see LICENSE.
This is a v0 skeleton, published early and openly rather than held back until every deferred module above is finished. Contributions that fill in a deferred piece (especially a clean, fully-generic worked example department loop) are welcome — open an issue or a PR.