feat(agent): finalize within time and turn budgets - #41
Merged
Merged
Conversation
Both remaining reward-0 tasks (caffe-cifar-10, mteb-leaderboard) passed with max_turns=100. mteb-leaderboard finished cleanly on turn 70; caffe-cifar-10 passed from artifacts on disk after hitting the deadline. The run removed the Harbor timeout overrides to stay leaderboard-compliant.
Headless runs can now be given a total wall-clock budget through --time-budget-seconds (Harbor kwarg time_budget_seconds). The loop tells the model how much time remains each turn, escalates to a stop-and-write warning inside a reserved finalization window, and stops with a new time_budget_exhausted outcome before the harness timeout kills the process. The budget is recorded in the journal and ATIF trajectory.
Experiment 2 adds a wall-clock budget, per-turn budget injection, and stop-on-deadline on top of Experiment 1. Both tasks passed and self-terminated cleanly with native trajectories. Also records that mutating the system prompt each turn broke prompt caching and raised cost about 9x.
Attribute the cache collapse to the per-turn system-prompt change using the DeepInfra before/after comparison, and record the concurrent provider-routing shift as a confounder in the structured results.
Rewrite the per-turn reminder into the most recent user message instead of the system prompt. A changing system prompt made the first tokens of every request differ, defeating the provider's prefix cache; appending to the tail keeps each request an extension of the previous one so the cached prefix survives.
- Recheck the wall-clock deadline before every tool in a reply, not only once before the batch, so a later tool cannot start after the budget expires. - Record each injected budget reminder as an input.injected event and project it into ATIF, so the behavior-changing input reaches the trajectory. - Bump the Event Journal schema to v4 for time_budget_exhausted and input.injected, following the response_truncated precedent, and update the protocol docs.
Experiment 3 moves the budget reminder to the conversation tail. Both tasks passed and the prefix cache recovered (input tokens 2.92M -> 0.109M, cost $0.856 -> $0.187), but mteb-leaderboard ran to the 100-turn cap instead of stopping on its own.
minixalpha
marked this pull request as ready for review
September 29, 2026 20:25
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Headless runs can finish useful work but continue until a reply cap or an external timeout prevents a normal final response. This change gives the model both remaining-time and remaining-reply guidance, asks it to save deliverables and finalize before either reserve is exhausted, and bounds blocking work so the agent can preserve its terminal record and trajectory.
--time-budget-secondsand the Harbor adapter'stime_budget_secondsargument. Keep the system prompt fixed and append runtime reminders to the conversation tail. Finalization starts with 10 replies remaining or within the larger of 180 seconds and 15% of the time budget.time_budget_exhausteddistinctly and limit subsequent cost reconciliation to 30 seconds before trajectory writing. Reject unsupported deadline environments.The PR includes an Unreleased changelog entry for the runtime budget and cleanup behavior, synchronized Chinese/English development notes, CLI and journal protocol documentation, and both Harbor READMEs. The root READMEs are unchanged. No Chinese research-note files are modified, so no research translation update is required.
Validation:
cdd65ebc99ed837f58467900db7892de82e68842.deepseek/deepseek-v4.1-flashonparasail/fp8, with fallback disabled: the repaired candidate passed 6/6 trials and completed normally in all six; the control passed 5/6 and completed normally in four. The registered pilot gate passed. Three repetitions per task are diagnostic, not a general reliability estimate.completed, no Harbor exceptions, supervisor deadlines, forced kills, or reconstructed trajectories. All 558 model requests were pinned and matched Parasail receipts. Caffe accuracy and PyTorch pipeline numerical checks remain task-solution failures; this PR does not claim to resolve them.Ordinary defaults remain 50 replies and 32768 output tokens, with no default wall-clock limit. The benchmark's 100/65536 settings and provider pin are experiment configuration. Historical score changes also involve model, routing, and budget differences, so they cannot be attributed solely to this implementation. Structured results and the full experiment report are included in the PR.