Skip to content

Repository files navigation

Open VS Code Agent

A coding agent designed for local open-weight models.

It runs entirely on a laptop with a 7B quantized model: no cloud API, no proprietary fallback. It is built and measured around the real constraints of local inference.

Status: public architecture edition Docs: CC BY 4.0 Code: MIT Measured: 24-task benchmark

Measured benchmark: Sprint 1E baseline vs Sprint 1G evidence-driven orchestration


Public repository scope. This is the public technical showcase of Open VS Code, a larger private development system. It publishes the architecture, the benchmark evidence, the evaluation method and a set of clean-room illustrative contracts. The private runtime, prompts, routing logic and benchmark fixtures are not included. More below.

Architecture

flowchart TB
  U([User intent]) --> M{Mode selection<br/>ASK · PLAN · DEBUG · REVIEW · BUILD}
  M --> C[Context assembly<br/>bounded budget · selective retrieval · project memory]
  C --> L[Local model<br/>Ollama · OpenAI-compatible wire]
  L --> N[Protocol normalization<br/>recover tool calls from imperfect output]
  N --> T{Tool selection<br/>schema-validated · mode-permitted}
  T -->|read-only| R[Repository tools]
  T -->|mutation| G[Durability gate<br/>write-ahead intent] --> X[Mutation<br/>base-hash checked · atomic]
  R --> O[Explicit tool output] --> C
  X --> V[Verification<br/>tests · typecheck · lint]
  V -->|pass| D([Result + evidence])
  V -->|fail| B[Bounded repair / rollback / escalate]
  B --> C
Loading

The model proposes, and a deterministic runtime decides what is allowed, what is recorded first, and what counts as done. → docs/architecture.md

Why local coding agents are hard

Frontier-model agents can lean on huge context windows, near-perfect tool calling and strong long-horizon reasoning. A 7B model at 4-bit quantization on a laptop has none of that:

  • Tool calls arrive malformed. They come in markdown fences, in ad-hoc tags, or with missing IDs. The runtime has to normalize them, not reject them.
  • "Done" gets declared too early. Small models tend to answer before they have looked at anything, so the runtime has to require evidence first.
  • Context is expensive. Every extra token costs latency on local hardware, and more context doesn't reliably buy accuracy (see the context experiment).
  • Loops and stalls are common. Turn budgets and non-progress detection are required, not optional.
  • Mistakes must be recoverable. A weaker model makes more of them, so every write has to be verified and reversible.

→ docs/local-model-constraints.md

Modes

Mode Purpose Side effects
ASK Explain code, answer questions about the repository None, read-only
PLAN Produce an implementation plan None, read-only
REVIEW Critique existing changes None, read-only
DEBUG Diagnose failures; may run known verification commands Verification only, human-gated
BUILD Make changes ("mutate") File mutations, human-gated, verified and reversible

The mode is a permission boundary, not a prompt style. Read-only modes can't mutate, whatever the model asks for.

Tool architecture

The tools are explicit and small, and every call is validated against a schema and the current mode before it runs:

Tool Class
read_file · list_directory · search_text · find_files Read-only, bounded output, ignore-aware
create_file · apply_patch Mutation: human-gated, base-hash checked, atomic, reversible
run_verification Runs discovered test / typecheck / lint scripts only; human-gated

MCP servers can add tools, subject to the same budgeting and risk mapping. → docs/tool-use.md

Context engineering

A local model gets less context, chosen more carefully: selective repository retrieval, compressed run state, explicit and bounded tool outputs, Markdown project memory (PROJECT.md, PLAN.md, ARCHITECTURE.md), and procedures that are disclosed progressively. → docs/context-management.md

Durability and mutation safety

intent recorded (write-ahead) → durability gate → mutation → verification → persist outcome → recover / escalate

If the intent can't be durably recorded, the mutation doesn't happen (fail closed). After a crash, recovery inspects what is actually on disk and resumes, re-verifies, asks a human, or stops. It never blindly replays a write. → docs/durability.md

Benchmark results

All numbers below are measured on the hardware listed further down, transcribed from raw result files. Design goals are labelled separately.

Sprint 1E vs 1G benchmark comparison

24-task read-only repository-understanding benchmark (20 dev + 4 held-out), Qwen2.5-Coder-7B-Instruct Q4_K_M via Ollama, 8k context:

Metric Sprint 1E baseline
2026-08-25
Sprint 1G evidence-driven
2026-08-26
End-to-end task success 37.5% (9/24) 75.0% (18/24)
Dev split 35.0% (7/20) 70.0% (14/20)
Held-out split 50.0% (2/4) 100.0% (4/4) ⚠︎ n=4
Tool-call syntax validity 90.9% (40/44) 93.2% (41/44)
Correct tool selection 79.2% (19/24) 95.8% (23/24)
Non-progress failures 4.2% (1/24) 8.3% (2/24)
Avg turns / task 2.63 3.29
Avg latency / task 6.1 s 9.4 s

What changed between the runs: Sprint 1G added evidence-driven orchestration. The runtime profiles the task, narrows the tool set, tracks which evidence has actually been gathered, and refuses final answers that aren't grounded in it. Success doubled. The cost was more turns and more latency, because the agent now reads before it answers.

Where it still fails: tasks about Markdown project rules (0/2) and some cross-file and control-flow reasoning. The remaining failures are mostly turn exhaustion and non-progress, not malformed tool calls.

Context-window experiment

Context window experiment

Context Pass rate Correct tool Avg turns Avg latency
2k 30% (3/10) 90% 2.8 5.3 s
8k 20% (2/10) 90% 2.6 3.7 s
32k 40% (4/10) 90% 3.1 7.1 s

On 10 tasks with the 1E agent, a 16× larger window changed the pass rate by only a couple of tasks, which is within noise for this sample size. It did roughly double latency relative to 8k. The takeaway is that careful context selection matters more than raw window size, not that one window size is best.

Runtime safety harness

A separate deterministic harness (30 scenarios, scripted model responses) exercises mutation, verification, repair and refusal of unsafe actions: 27 passed and 3 were correctly rejected as unsafe. This validates the runtime's mechanics. It is not a measure of model capability.

→ Methodology, limitations and raw aggregates: benchmarks/

Example trace

A bounded file mutation in BUILD mode, fully synthetic:

[BUILD] "Make slugify() handle accented characters"
 1 search_text   "function slugify"          → src/text/slugify.ts:3
 2 read_file     src/text/slugify.ts          (38 lines)
 3 read_file     tests/slugify.test.ts        (1 failing case identified)
 4 apply_patch   proposed · +2 −1 · base sha256:4be1…  ⏸ awaiting approval
   ✓ approved → intent journaled → base hash matches → atomic write
 5 run_verification  npm test                 ⏸ approval → 14 passed
 ✔ done · evidence: 2 files read, 1 patch applied, tests green

More: examples/synthetic-traces/

Hardware and model configuration

Machine Apple M3 Pro (11-core CPU), 18 GB unified memory
Inference Ollama v0.32.15, OpenAI-compatible endpoint
Model Qwen2.5-Coder-7B-Instruct, Q4_K_M (~4.7 GB)
Context 8k by default; 2k and 32k tested
Language TypeScript on Node.js ≥ 22; editor-independent core

The gateway speaks the OpenAI-compatible protocol, so llama.cpp and vLLM servers are also supported targets. The numbers above were measured only on the configuration in this table.

Project status

Area Status
Editor-independent agent core, deterministic tool loop Implemented
Local model gateway + protocol normalization Implemented
Read-only repository tools, context compiler, project memory Implemented
Evidence-driven orchestration Implemented and measured (Sprint 1G)
Safe mutation, controlled verification, bounded repair Implemented; validated by the scripted safety harness
Durable checkpoints and crash recovery Implemented; deterministic recovery test suite
Bounded MCP support, skills, interactive CLI Implemented
VS Code / Code-OSS extension integration Planned
Live-model benchmark of BUILD / DEBUG modes Not yet measured; the published live numbers cover read-only tasks only

Public repository scope

This repository is a public technical showcase and a selected, clean-room implementation surface of a larger private system.

Included: architecture and design documentation written for this edition, benchmark aggregates with methodology and limitations, illustrative TypeScript contracts (type-only), and synthetic traces over made-up repositories.

Intentionally not included: the agent runtime, all prompts and templates, routing, profiling and tool-narrowing rules, context-budget and compression parameters, the mutation and recovery implementation, benchmark task prompts, fixture repositories and expected answers, and configuration files.

This repository has its own history. It was not forked, mirrored or filtered from the private one.


© Raul Mermans · docs CC BY 4.0, code MIT (see LICENSE) · SECURITY.md

About

Local coding agent designed and benchmarked for open-weight models, with explicit tools, verification and recoverable mutation.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages