Skip to content

feat(agent): finalize within time and turn budgets - #41

Merged
minixalpha merged 12 commits into
mainfrom
experiment/max-turns-and-time-budget
Sep 30, 2026
Merged

minixalpha merged 12 commits into
mainfrom
experiment/max-turns-and-time-budget

Conversation

@minixalpha

@minixalpha minixalpha commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Headless runs can finish useful work but continue until a reply cap or an external timeout prevents a normal final response. This change gives the model both remaining-time and remaining-reply guidance, asks it to save deliverables and finalize before either reserve is exhausted, and bounds blocking work so the agent can preserve its terminal record and trajectory.

  • Add --time-budget-seconds and the Harbor adapter's time_budget_seconds argument. Keep the system prompt fixed and append runtime reminders to the conversation tail. Finalization starts with 10 replies remaining or within the larger of 180 seconds and 15% of the time budget.
  • Enforce wall-clock deadlines for model calls, individual tools, and retry eligibility on POSIX main threads. Record time_budget_exhausted distinctly and limit subsequent cost reconciliation to 30 seconds before trajectory writing. Reject unsupported deadline environments.
  • Bound bash cleanup after timeout or interruption: terminate the command's process group and other session members on Linux, close inherited output pipes, and bound direct-child reaping. Normally completed commands preserve background services.
  • Record effective budgets and injected reminders in Event Journal schema v4 and ATIF, retaining replay support for v1-v3 journals.

The PR includes an Unreleased changelog entry for the runtime budget and cleanup behavior, synchronized Chinese/English development notes, CLI and journal protocol documentation, and both Harbor READMEs. The root READMEs are unchanged. No Chinese research-note files are modified, so no research translation update is required.

Validation:

  • 354 main-package and Harbor compatibility tests passed in the pinned benchmark environment, including real-process cleanup regressions. GitHub CI passed on Python 3.13 and 3.14 at cdd65ebc99ed837f58467900db7892de82e68842.
  • Fixed-provider targeted comparison with deepseek/deepseek-v4.1-flash on parasail/fp8, with fallback disabled: the repaired candidate passed 6/6 trials and completed normally in all six; the control passed 5/6 and completed normally in four. The registered pilot gate passed. Three repetitions per task are diagnostic, not a general reliability estimate.
  • The subsequent pinned pilot20 scored 18/20, with all 20 runs ending as native completed, no Harbor exceptions, supervisor deadlines, forced kills, or reconstructed trajectories. All 558 model requests were pinned and matched Parasail receipts. Caffe accuracy and PyTorch pipeline numerical checks remain task-solution failures; this PR does not claim to resolve them.

Ordinary defaults remain 50 replies and 32768 output tokens, with no default wall-clock limit. The benchmark's 100/65536 settings and provider pin are experiment configuration. Historical score changes also involve model, routing, and budget differences, so they cannot be attributed solely to this implementation. Structured results and the full experiment report are included in the PR.

Both remaining reward-0 tasks (caffe-cifar-10, mteb-leaderboard) passed with
max_turns=100. mteb-leaderboard finished cleanly on turn 70; caffe-cifar-10
passed from artifacts on disk after hitting the deadline. The run removed the
Harbor timeout overrides to stay leaderboard-compliant.
Headless runs can now be given a total wall-clock budget through
--time-budget-seconds (Harbor kwarg time_budget_seconds). The loop tells the
model how much time remains each turn, escalates to a stop-and-write warning
inside a reserved finalization window, and stops with a new
time_budget_exhausted outcome before the harness timeout kills the process.
The budget is recorded in the journal and ATIF trajectory.
Experiment 2 adds a wall-clock budget, per-turn budget injection, and
stop-on-deadline on top of Experiment 1. Both tasks passed and self-terminated
cleanly with native trajectories. Also records that mutating the system prompt
each turn broke prompt caching and raised cost about 9x.
Attribute the cache collapse to the per-turn system-prompt change using the
DeepInfra before/after comparison, and record the concurrent provider-routing
shift as a confounder in the structured results.
Rewrite the per-turn reminder into the most recent user message instead of the
system prompt. A changing system prompt made the first tokens of every request
differ, defeating the provider's prefix cache; appending to the tail keeps each
request an extension of the previous one so the cached prefix survives.
- Recheck the wall-clock deadline before every tool in a reply, not only once
  before the batch, so a later tool cannot start after the budget expires.
- Record each injected budget reminder as an input.injected event and project
  it into ATIF, so the behavior-changing input reaches the trajectory.
- Bump the Event Journal schema to v4 for time_budget_exhausted and
  input.injected, following the response_truncated precedent, and update the
  protocol docs.
Experiment 3 moves the budget reminder to the conversation tail. Both tasks
passed and the prefix cache recovered (input tokens 2.92M -> 0.109M, cost
$0.856 -> $0.187), but mteb-leaderboard ran to the 100-turn cap instead of
stopping on its own.
@minixalpha
minixalpha marked this pull request as ready for review September 29, 2026 20:25
@minixalpha
minixalpha merged commit e75cacb into main Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant