A coding agent designed for local open-weight models.
It runs entirely on a laptop with a 7B quantized model: no cloud API, no proprietary fallback. It is built and measured around the real constraints of local inference.
Public repository scope. This is the public technical showcase of Open VS Code, a larger private development system. It publishes the architecture, the benchmark evidence, the evaluation method and a set of clean-room illustrative contracts. The private runtime, prompts, routing logic and benchmark fixtures are not included. More below.
flowchart TB
U([User intent]) --> M{Mode selection<br/>ASK · PLAN · DEBUG · REVIEW · BUILD}
M --> C[Context assembly<br/>bounded budget · selective retrieval · project memory]
C --> L[Local model<br/>Ollama · OpenAI-compatible wire]
L --> N[Protocol normalization<br/>recover tool calls from imperfect output]
N --> T{Tool selection<br/>schema-validated · mode-permitted}
T -->|read-only| R[Repository tools]
T -->|mutation| G[Durability gate<br/>write-ahead intent] --> X[Mutation<br/>base-hash checked · atomic]
R --> O[Explicit tool output] --> C
X --> V[Verification<br/>tests · typecheck · lint]
V -->|pass| D([Result + evidence])
V -->|fail| B[Bounded repair / rollback / escalate]
B --> C
The model proposes, and a deterministic runtime decides what is allowed, what is recorded first, and what counts as done. → docs/architecture.md
Frontier-model agents can lean on huge context windows, near-perfect tool calling and strong long-horizon reasoning. A 7B model at 4-bit quantization on a laptop has none of that:
- Tool calls arrive malformed. They come in markdown fences, in ad-hoc tags, or with missing IDs. The runtime has to normalize them, not reject them.
- "Done" gets declared too early. Small models tend to answer before they have looked at anything, so the runtime has to require evidence first.
- Context is expensive. Every extra token costs latency on local hardware, and more context doesn't reliably buy accuracy (see the context experiment).
- Loops and stalls are common. Turn budgets and non-progress detection are required, not optional.
- Mistakes must be recoverable. A weaker model makes more of them, so every write has to be verified and reversible.
→ docs/local-model-constraints.md
| Mode | Purpose | Side effects |
|---|---|---|
| ASK | Explain code, answer questions about the repository | None, read-only |
| PLAN | Produce an implementation plan | None, read-only |
| REVIEW | Critique existing changes | None, read-only |
| DEBUG | Diagnose failures; may run known verification commands | Verification only, human-gated |
| BUILD | Make changes ("mutate") | File mutations, human-gated, verified and reversible |
The mode is a permission boundary, not a prompt style. Read-only modes can't mutate, whatever the model asks for.
The tools are explicit and small, and every call is validated against a schema and the current mode before it runs:
| Tool | Class |
|---|---|
read_file · list_directory · search_text · find_files |
Read-only, bounded output, ignore-aware |
create_file · apply_patch |
Mutation: human-gated, base-hash checked, atomic, reversible |
run_verification |
Runs discovered test / typecheck / lint scripts only; human-gated |
MCP servers can add tools, subject to the same budgeting and risk mapping. → docs/tool-use.md
A local model gets less context, chosen more carefully: selective repository retrieval, compressed run state, explicit and bounded tool outputs, Markdown project memory (PROJECT.md, PLAN.md, ARCHITECTURE.md), and procedures that are disclosed progressively. → docs/context-management.md
intent recorded (write-ahead) → durability gate → mutation → verification → persist outcome → recover / escalate
If the intent can't be durably recorded, the mutation doesn't happen (fail closed). After a crash, recovery inspects what is actually on disk and resumes, re-verifies, asks a human, or stops. It never blindly replays a write. → docs/durability.md
All numbers below are measured on the hardware listed further down, transcribed from raw result files. Design goals are labelled separately.
24-task read-only repository-understanding benchmark (20 dev + 4 held-out), Qwen2.5-Coder-7B-Instruct Q4_K_M via Ollama, 8k context:
| Metric | Sprint 1E baseline 2026-08-25 |
Sprint 1G evidence-driven 2026-08-26 |
|---|---|---|
| End-to-end task success | 37.5% (9/24) | 75.0% (18/24) |
| Dev split | 35.0% (7/20) | 70.0% (14/20) |
| Held-out split | 50.0% (2/4) | 100.0% (4/4) ⚠︎ n=4 |
| Tool-call syntax validity | 90.9% (40/44) | 93.2% (41/44) |
| Correct tool selection | 79.2% (19/24) | 95.8% (23/24) |
| Non-progress failures | 4.2% (1/24) | 8.3% (2/24) |
| Avg turns / task | 2.63 | 3.29 |
| Avg latency / task | 6.1 s | 9.4 s |
What changed between the runs: Sprint 1G added evidence-driven orchestration. The runtime profiles the task, narrows the tool set, tracks which evidence has actually been gathered, and refuses final answers that aren't grounded in it. Success doubled. The cost was more turns and more latency, because the agent now reads before it answers.
Where it still fails: tasks about Markdown project rules (0/2) and some cross-file and control-flow reasoning. The remaining failures are mostly turn exhaustion and non-progress, not malformed tool calls.
| Context | Pass rate | Correct tool | Avg turns | Avg latency |
|---|---|---|---|---|
| 2k | 30% (3/10) | 90% | 2.8 | 5.3 s |
| 8k | 20% (2/10) | 90% | 2.6 | 3.7 s |
| 32k | 40% (4/10) | 90% | 3.1 | 7.1 s |
On 10 tasks with the 1E agent, a 16× larger window changed the pass rate by only a couple of tasks, which is within noise for this sample size. It did roughly double latency relative to 8k. The takeaway is that careful context selection matters more than raw window size, not that one window size is best.
A separate deterministic harness (30 scenarios, scripted model responses) exercises mutation, verification, repair and refusal of unsafe actions: 27 passed and 3 were correctly rejected as unsafe. This validates the runtime's mechanics. It is not a measure of model capability.
→ Methodology, limitations and raw aggregates: benchmarks/
A bounded file mutation in BUILD mode, fully synthetic:
[BUILD] "Make slugify() handle accented characters"
1 search_text "function slugify" → src/text/slugify.ts:3
2 read_file src/text/slugify.ts (38 lines)
3 read_file tests/slugify.test.ts (1 failing case identified)
4 apply_patch proposed · +2 −1 · base sha256:4be1… ⏸ awaiting approval
✓ approved → intent journaled → base hash matches → atomic write
5 run_verification npm test ⏸ approval → 14 passed
✔ done · evidence: 2 files read, 1 patch applied, tests green
More: examples/synthetic-traces/
| Machine | Apple M3 Pro (11-core CPU), 18 GB unified memory |
| Inference | Ollama v0.32.15, OpenAI-compatible endpoint |
| Model | Qwen2.5-Coder-7B-Instruct, Q4_K_M (~4.7 GB) |
| Context | 8k by default; 2k and 32k tested |
| Language | TypeScript on Node.js ≥ 22; editor-independent core |
The gateway speaks the OpenAI-compatible protocol, so llama.cpp and vLLM servers are also supported targets. The numbers above were measured only on the configuration in this table.
| Area | Status |
|---|---|
| Editor-independent agent core, deterministic tool loop | Implemented |
| Local model gateway + protocol normalization | Implemented |
| Read-only repository tools, context compiler, project memory | Implemented |
| Evidence-driven orchestration | Implemented and measured (Sprint 1G) |
| Safe mutation, controlled verification, bounded repair | Implemented; validated by the scripted safety harness |
| Durable checkpoints and crash recovery | Implemented; deterministic recovery test suite |
| Bounded MCP support, skills, interactive CLI | Implemented |
| VS Code / Code-OSS extension integration | Planned |
| Live-model benchmark of BUILD / DEBUG modes | Not yet measured; the published live numbers cover read-only tasks only |
This repository is a public technical showcase and a selected, clean-room implementation surface of a larger private system.
Included: architecture and design documentation written for this edition, benchmark aggregates with methodology and limitations, illustrative TypeScript contracts (type-only), and synthetic traces over made-up repositories.
Intentionally not included: the agent runtime, all prompts and templates, routing, profiling and tool-narrowing rules, context-budget and compression parameters, the mutation and recovery implementation, benchmark task prompts, fixture repositories and expected answers, and configuration files.
This repository has its own history. It was not forked, mirrored or filtered from the private one.
© Raul Mermans · docs CC BY 4.0, code MIT (see LICENSE) · SECURITY.md