Skip to content

reconcile() marks all in-flight executions failed on every load, including hot reloads #39

Description

@Fishsb

Summary

ExecutionService.reconcile() runs on every plugin load and unconditionally marks every outcome: 'running' execution as failed / "interrupted by host restart", handing its task back to todo:

async reconcile(): Promise<void> {
  await this.deps.store.mutate('execution-recorded', (ledger) => {
    for (const task of ledger.tasks) {
      for (const execution of task.executions) {
        if (execution.outcome === 'running') {
          execution.outcome = 'failed'
          execution.error = 'interrupted by host restart'
          // …task back to todo, claim released

Its doc comment states the assumption: "executions left running by the previous process can never settle here (their settlement watchers died with it)". That holds for a real host restart — but not for a hot reload, which is the common case while developing or updating the plugin.

Why it is wrong

A hot reload replaces the plugin fiber but keeps the host process and its live agents. Measured on a live host: pid and /proc/<pid>/stat start-time are both unchanged across a reload, and the previously-started whenIdle() settlement watchers are still pending. So:

  • the settlements are not dead — they fire later and write their result;
  • meanwhile reconcile() has already recorded those runs as failed and moved their cards to todo;
  • and because the run was released from runs, the late settlement can also produce a task whose card says todo while its execution says succeeded.

Every reload therefore kills all in-flight executions. I hit this repeatedly while reloading a patched build: 4 live runs (3 belonging to other sessions) were marked failed by a single reload, each leaving a 失败 record and a system comment on the card. One of them could not be recovered because its session had already ended.

Suggested direction

Ask liveness instead of assuming a restart: an execution whose owning session still resolves from the agent registry has a live watcher and must be left running. Only runs whose session is genuinely gone can never settle and should be failed. A successor generation should then adopt the surviving runs (re-attach a whenIdle watcher), otherwise the records stay running forever.

Note the two neighbouring pieces if you take this on:

  • adoption needs the run's PreparedMirror for evidence collection. prepared is currently memory-only, so an adopted worktree run settles with no commits / headCommit / changedFiles — reconstructing it from the persisted repos / worktreePath / baseCommit fields works.
  • two store instances then share one file, which is what fix(store): never write the ledger from a snapshot older than the file #38 (the stale-write rollback) is about; that PR is worth landing first, since adoption widens the window in which two generations coexist.

Reproduction

  1. Have at least one task in in_progress with a running execution (a live session with a long turn).
  2. Hot-reload the plugin (dev_reload_package dsh-taskboard).
  3. The execution is now failed / interrupted by host restart, the card is back in todo, and a [系统] 执行失败 comment appeared — while the agent session is still running.

Happy to send a PR for this if useful; I have a working implementation and tests.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions