I'm Yash. I build open-source infrastructure that sits between an AI agent and the real world. It records what the agent did, checks what it claims, and pins down what it's allowed to call. Everything is local-first and model-agnostic, built on boring formats you can inspect.
| Project | What it proves |
|---|---|
| RunLedger | What the agent actually did. Tool calls, decisions and retries are hash-chained, so any edit, deletion or reordering breaks verification. |
| BrowserProof | That the browser really ended up where the agent claims. Explicit assertions, evidence hashes, CI exit codes. |
| ActionMesh | What the agent is allowed to call. One typed contract per capability, served over MCP-style, HTTP and CLI. |
| Verify | That the output is right. Checks DOCX/XLSX/PPTX/PDF files, repos and reversible actions without trusting the generator. |
| Agent Compat Lab | That one repo tells Codex, Claude Code, Gemini CLI and OpenCode the same thing. Outputs SARIF. |
| Agent Reliability Lab | Twelve deterministic tools for hardening agent infrastructure: MCP chaos, contract fuzzing, context firewalls, replay, evals and more. |
npm install github:yashkhou/runledger && npx runledger verify- Glyph: an experimental design language with memory, so agents can reason about structure instead of pixels.
- Commander Plus: a local-first agent workstation with MCP workspaces, reusable skills, browser control and persistent project context.
- OpenRetention: self-hosted customer-success software with explainable health scores and revenue-at-risk prioritisation.
- Opened Stellar-agentic #429 to repair repository-specific README links.
- Reviewed pydantic-ai #8969, a regression fix preventing shared
StructuredDictschema metadata from leaking between output types. - Added current-main implementation analysis to MCP Python SDK #1933 around stdio ownership and process-stream lifecycle.
If one of these saves you a bad night, a ⭐ helps the next person find it.



