You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
reconcile() marks all in-flight executions failed on every load, including hot reloads #39
ExecutionService.reconcile() runs on every plugin load and unconditionally marks every outcome: 'running' execution as failed / "interrupted by host restart", handing its task back to todo:
asyncreconcile(): Promise<void>{awaitthis.deps.store.mutate('execution-recorded',(ledger)=>{for(consttaskofledger.tasks){for(constexecutionoftask.executions){if(execution.outcome==='running'){execution.outcome='failed'execution.error='interrupted by host restart'// …task back to todo, claim released
Its doc comment states the assumption: "executions left running by the previous process can never settle here (their settlement watchers died with it)". That holds for a real host restart — but not for a hot reload, which is the common case while developing or updating the plugin.
Why it is wrong
A hot reload replaces the plugin fiber but keeps the host process and its live agents. Measured on a live host: pid and /proc/<pid>/stat start-time are both unchanged across a reload, and the previously-started whenIdle() settlement watchers are still pending. So:
the settlements are not dead — they fire later and write their result;
meanwhile reconcile() has already recorded those runs as failed and moved their cards to todo;
and because the run was released from runs, the late settlement can also produce a task whose card says todo while its execution says succeeded.
Every reload therefore kills all in-flight executions. I hit this repeatedly while reloading a patched build: 4 live runs (3 belonging to other sessions) were marked failed by a single reload, each leaving a 失败 record and a system comment on the card. One of them could not be recovered because its session had already ended.
Suggested direction
Ask liveness instead of assuming a restart: an execution whose owning session still resolves from the agent registry has a live watcher and must be left running. Only runs whose session is genuinely gone can never settle and should be failed. A successor generation should then adopt the surviving runs (re-attach a whenIdle watcher), otherwise the records stay running forever.
Note the two neighbouring pieces if you take this on:
adoption needs the run's PreparedMirror for evidence collection. prepared is currently memory-only, so an adopted worktree run settles with no commits / headCommit / changedFiles — reconstructing it from the persisted repos / worktreePath / baseCommit fields works.
Have at least one task in in_progress with a running execution (a live session with a long turn).
Hot-reload the plugin (dev_reload_package dsh-taskboard).
The execution is now failed / interrupted by host restart, the card is back in todo, and a [系统] 执行失败 comment appeared — while the agent session is still running.
Happy to send a PR for this if useful; I have a working implementation and tests.
Summary
ExecutionService.reconcile()runs on every plugin load and unconditionally marks everyoutcome: 'running'execution asfailed/"interrupted by host restart", handing its task back totodo:Its doc comment states the assumption: "executions left
runningby the previous process can never settle here (their settlement watchers died with it)". That holds for a real host restart — but not for a hot reload, which is the common case while developing or updating the plugin.Why it is wrong
A hot reload replaces the plugin fiber but keeps the host process and its live agents. Measured on a live host: pid and
/proc/<pid>/statstart-time are both unchanged across a reload, and the previously-startedwhenIdle()settlement watchers are still pending. So:reconcile()has already recorded those runs asfailedand moved their cards totodo;runs, the late settlement can also produce a task whose card saystodowhile its execution sayssucceeded.Every reload therefore kills all in-flight executions. I hit this repeatedly while reloading a patched build: 4 live runs (3 belonging to other sessions) were marked failed by a single reload, each leaving a
失败record and a system comment on the card. One of them could not be recovered because its session had already ended.Suggested direction
Ask liveness instead of assuming a restart: an execution whose owning session still resolves from the agent registry has a live watcher and must be left
running. Only runs whose session is genuinely gone can never settle and should be failed. A successor generation should then adopt the surviving runs (re-attach awhenIdlewatcher), otherwise the records stayrunningforever.Note the two neighbouring pieces if you take this on:
PreparedMirrorfor evidence collection.preparedis currently memory-only, so an adopted worktree run settles with nocommits/headCommit/changedFiles— reconstructing it from the persistedrepos/worktreePath/baseCommitfields works.Reproduction
in_progresswith a running execution (a live session with a long turn).dev_reload_package dsh-taskboard).failed/interrupted by host restart, the card is back intodo, and a[系统] 执行失败comment appeared — while the agent session is still running.Happy to send a PR for this if useful; I have a working implementation and tests.