Skip to content

feat(one-agent): make one job = one one call a first-class RAS rule - #10

Merged
yaonyan merged 9 commits into
mainfrom
feat/prompt-batch-first-class
Sep 18, 2026
Merged

yaonyan merged 9 commits into
mainfrom
feat/prompt-batch-first-class

Conversation

@yaonyan

@yaonyan yaonyan commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Addresses the gap reported in #8 (comment): the resident prompt told models to batch but the model still opened multiple one calls per job, losing the control node where reason() belongs. The old wording ("One bounded action phase = one RAS call") was buried in Directives and abstract; this PR rewrites the mental model using the round-trip cost that models already understand from ReAct training data, and backs it with a stricter control-node eval.

Changes

  • src/prompts.ts — rewrite the resident mental model:
    • Frame one as a working session (sandbox), contrasted with ReAct's reason-act-observe round-trip tax; every one call appends raw output to the shared conversation and consumes attention budget.
    • "One job = one one call" is a first-class planning default: batch all currently known related work into one bounded call. Starting a new call is allowed only when the previous one failed, timed out, or returned evidence that reveals a genuinely new target or dependency.
    • reason() is judgment-only at genuinely uncertain convergence points (targeting, branching, retry-vs-escalate, classification, synthesis); deterministic policies (thresholds, exit codes, rule-based transforms) stay in code/shell.
    • Inputs contract aligned across bash/python/typescript mode deltas: only values piped into act <tool> - must be JSON objects matching the tool schema; strings, arrays, and primitives are allowed elsewhere.
  • one-runner-bash.md / one-runner-python.md — same one-job-one-call rule and inputs contract for on-demand runner docs.
  • src/ras/*.ts — inputs schema descriptions match the relaxed contract for bash, python, and typescript.
  • evals/agent_eval.ts:
    • Evidence-stream counting now strips comments and quoted string literals first, so act bash '{"command":"git log"}' counts one stream (not two) and cat/git mentions in prompts no longer inflate the batch score.
    • converge requires the exact normalized file set plus a non-empty reason (no more substring superset pass).
  • evals/datasets/ras-control-node.json + evals/fixtures/changed-files.txt — add REPL-style cn-repl-converge probe forcing 2–3 evidence streams to converge in one call (expectBatched: true).
  • tests/prompts.test.ts — assert the new mental-model and inputs-contract wording across all three modes.
  • vitest.config.ts — retry is CI-only so local flaky regressions surface immediately.

Validation

  • pnpm --filter=@one-agent/agent test → 108/108 pass (Test step runs in CI).
  • tsc --noEmit clean.
  • Eval run (deepseek-v4-flash):
    • ras-control-node: model now batches the initial evidence read (one call reads all 3 fixtures) vs. the previously observed 5 separate round trips; the strict eval still flags extra follow-up calls, which is the regression signal we want to keep.
    • directanswer gate unchanged vs. baseline (no regression).

Rewrite the resident mental model around the ReAct round-trip cost that
models already know: every `one` call appends raw output to the shared
conversation and consumes attention budget for later steps, so splitting
a job across several calls leaks intermediate state into the main
context. One job = one `one` call is now a first-class rule with
failure-mode phrasing, batching is unconditional, and reason() is
explicitly judgment-only at convergence points.

Eval: count bash shell evidence streams (cat/sed/git/...) toward the
batched signal, add a REPL-style converge probe that forces 2-3 evidence
streams to converge in one call, and add a semantic check for it.

Runner docs (bash/python) and prompt tests updated to match.
A live session showed the model burning a dozen round trips creating a
file: it invented a __ONE_INPUT__ placeholder (written literally), mixed
positional JSON with `-` in act (Unknown argument: -), and tried jq
--arg/python3 to build JSON — both unavailable in the minimal sandbox.

- parseActArgs: when `-` is combined with positional JSON, fail with the
  correct usage (one-input <key> | act <tool> -) instead of a bare
  "Unknown argument: -".
- prompts.ts bash delta + one-runner-bash.md: state there is no
  placeholder substitution, and document sandbox limits (jq has no
  --arg/--slurpfile; python3 is unavailable for arg construction).
- tests: stdin-JSON path via `act <tool> -`, the mixing rejection, and
  the bash prompt wording.
A live edit session failed because the caller passed a JSON-encoded
string for inputs.edit_args; one-input JSON.stringifies it again, so act
parsed it back into a string and edit saw no path field. State the
contract where the model reads it: inputs schema descriptions (bash,
python, typescript), the bash runner doc, and the bash system prompt.
When a caller passes a JSON-encoded string for inputs.edit_data (instead
of an object), one-input JSON.stringifies it again and act parses it back
into a string; edit then destructured undefined oldText and crashed with
"Cannot read properties of undefined (reading 'split')", which led the
model to wrongly conclude edit does not accept stdin.

Validate path/oldText/newText at the edit entry and explain the contract:
pass the object, not a JSON-encoded string.
- countEvidenceStreams strips comments and quoted strings so an
  `act bash '{"command":"git log"}'` counts one evidence stream, not two,
  and cat/git mentions in prompts no longer inflate the batch score
- converge now requires the exact normalized file set plus a non-empty
  reason instead of a substring superset match
- one job = one `one` call is a planning default with explicit recovery
  exceptions (previous call failed, timed out, or revealed a genuinely
  new target or dependency)
- inputs contract: only values piped into `act <tool> -` must be JSON
  objects matching the tool schema; strings, arrays, and primitives are
  allowed elsewhere, aligned across bash, python, and typescript prompts
- vitest retry is CI-only so local flaky regressions surface immediately
@yaonyan

yaonyan commented Aug 16, 2026

Copy link
Copy Markdown
Collaborator Author

Review fixes (commit 41e8686)

Addressed all Bugbot findings:

  • [P1] Evidence-stream double counting — countEvidenceStreams now strips comments and quoted string literals before counting (stripCodeNoise), so act bash '{"command":"git log"}' counts one stream instead of two, and cat/git mentions in prompts or comments no longer inflate the batch score.
  • [P1] converge false positives — now requires the exact normalized file set (got.size === want.size + full match, no substring superset) and a non-empty reason, so returning all files or not-src/b.test.ts.bak no longer passes.
  • [P2] "One job = one call" as an absolute constraint — reworded to a planning default with explicit recovery exceptions: start a new one call only when the previous one failed, timed out, or returned evidence revealing a genuinely new target or dependency. Synced across prompts.ts and both runner docs.
  • [P2] inputs values must be objects — narrowed to "only values piped into act <tool> - must be JSON objects matching the tool's argument schema; strings, arrays, and primitives are allowed elsewhere". Applied consistently to the bash/python/typescript mode deltas, the src/ras/*.ts schema descriptions, and one-runner-bash.md.
  • Non-blocking: global retry: 1 — now process.env.CI ? 1 : 0 so local runs surface flaky regressions immediately.

Validation: pnpm --filter=@one-agent/agent test 108/108 pass, tsc --noEmit clean, CI green on this head (build job includes build/typecheck/Test).

The edit tool is shared across modes; the one-input guidance in the
validation error only applies to Bash. Keep the generic actionable hint.
The __ONE_INPUT__ warning adds noise; the one-input pipe guidance is
enough. Keeps the pass-values-with-one-input directive in both the
resident bash delta and the runner doc.
The Pyodide/Deno runtime names add no decision value; mode deltas now
just say what one.code is.
@yaonyan
yaonyan merged commit 84f4031 into main Sep 18, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant