Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
9d46999
Remove the GPU resource loan subsystem
praveenperera Oct 3, 2026
747fc47
Add run process cleanup and a preempted outcome
praveenperera Oct 3, 2026
2cc70af
Fix a cancel race in the follow-up blocker test
praveenperera Oct 3, 2026
5b8f3ea
Add the GPU queue domain, store, and scheduler
praveenperera Oct 3, 2026
f377f64
Run the GPU queue with a per-machine actor
praveenperera Oct 3, 2026
94f6c77
Report why a sender test daemon was not ready
praveenperera Oct 3, 2026
f180295
Expose the GPU queue through routes, events, and CLI
praveenperera Oct 3, 2026
54a9686
Open a job directory to other users only for containers
praveenperera Oct 3, 2026
1ca81c3
Keep job submit off the TCP listener
praveenperera Oct 3, 2026
86ed2b6
Document the GPU priority queue
praveenperera Oct 3, 2026
efe8891
Retry a sender test daemon on a fresh port
praveenperera Oct 3, 2026
97f312d
Show the GPU queue in the web dashboard
praveenperera Oct 3, 2026
190bdfd
Fix races and gaps found in the queue review
praveenperera Oct 3, 2026
2b34a31
Report Linux zombies as exited processes
praveenperera Oct 3, 2026
e665799
Make the platform binary cleanup test host-independent
praveenperera Oct 3, 2026
e6ddc54
Make two tests pass on CI hosts
praveenperera Oct 3, 2026
a12610f
Preempt for any waiting job that can use a GPU
praveenperera Oct 3, 2026
2affb97
Test an upgrade from a database a real release wrote
praveenperera Oct 3, 2026
ebefdff
Let cleanup finish on a busy machine
praveenperera Oct 3, 2026
a2803d3
Hold cleanup on unreadable processes a run could own
praveenperera Oct 3, 2026
6c51c67
Test upgrades from databases every release wrote
praveenperera Oct 3, 2026
eb0cd70
Let a vanished process delay cleanup confirmation
praveenperera Oct 3, 2026
75d6d26
Refactor the GPU queue code for clarity
praveenperera Oct 3, 2026
516ee46
Simplify the GPU queue code and drop restating tests
praveenperera Oct 3, 2026
fab850d
Refactor individual files for readability
praveenperera Oct 3, 2026
378cfae
Reshape the data model and start a fresh database
praveenperera Oct 3, 2026
5b90a8c
Split large modules and remove duplicated code
praveenperera Oct 3, 2026
8e3d046
Fix bugs found during the refactor passes
praveenperera Oct 4, 2026
0a405f4
Save the T3 send intent atomically
praveenperera Oct 4, 2026
540d2ff
Move the delivery tests to the module layout
praveenperera Oct 4, 2026
0c7b003
Expect quoted agent paths in the systemd dry run test
praveenperera Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .agents/skills/homebased/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,11 @@ description: Run long, unattended agent CLIs and general task commands through t
- Never run `codex queue` yourself, never write to a Claude Code messaging socket yourself, and never tell a worker to do either. Delivery belongs to `homebased`.
- A Claude Code session is a root agent like a Codex thread. Its thread id is `$CLAUDE_CODE_SESSION_ID`. The daemon sends events for that id to the live session through its messaging socket, and sends events for any other id through `codex queue`. When a T3 Code thread owns the session or Codex thread, the daemon starts the turn through T3 instead, so the thread shows it; [messages.md](references/messages.md) has the exact rules.
- Submit through the daemon on the machine that owns the Codex thread. To run the child elsewhere, enable Fleet and set `machine` in the JSON spec. The submitting machine remains the origin and sends callbacks to the original thread; the selected Fleet machine executes the child.
- Send all GPU work through the resource queue. Read [resource-queue.md](references/resource-queue.md) for priority, checkpoints, job events, and cleanup.
- Put task fields in the JSON spec and submit with `homebased task submit --spec <file|->`. Use `--request-id <uuid>` when a caller needs a stable retry identity.
- Always set `name` to a short goal label. Do not name the task after the agent or the CLI.
- Always pass `--json` on data commands and parse the result. Every JSON object carries `api_version: 1`.
- Event delivery is at-least-once. For new events, deduplicate by `(task, seq)`; for a legacy event without `seq`, use `(task, event)`.
- For GPU resource work, read [resource-loans.md](references/resource-loans.md). Check the exact pending actions at start, after compaction, and before an independent background launch. A delivered notice is not completion, and a notice whose action is absent from a complete pending result is stale. An unavailable authority is not an empty action list.
- Event delivery is at-least-once. Deduplicate by `(task, seq)`.
- Use `homebased message send` for a direct message to a Claude Code session or Codex thread. Address a session by its id or by a task whose origin it is, never by its display name, which changes each time the session restarts. Read [messages.md](references/messages.md) for destination, source, and retry rules.
- `message send --task` targets the task's origin thread, not its worker. Use `message send --worker <task>` to give a running Claude worker new instructions; it reads them at its next turn boundary. Use `homebased task followup` to resume a terminal Codex worker with new information. Run follow-up on the task's origin or execution machine. Only one follow-up can resume a thread at a time; wait for the active task's event after `resume_thread_busy`.
- Do not poll a running task in a loop. Submit, tell the user the task id, end the turn, and wait for events. Inspect on demand only.
Expand All @@ -37,12 +37,12 @@ Pick the first row that matches, then read only that file.
| --- | --- |
| `HOMEBASED_TASK_ID` is set in this session's environment | [worker.md](references/worker.md). You are the worker, not the orchestrator. |
| A message starting with `HOMEBASED_EVENT ` arrived | [events.md](references/events.md) |
| GPU work, resource queues, job specs, priority, preemption, or Attention | [resource-queue.md](references/resource-queue.md) |
| Starting background work, following up a finished task, chaining work with `after`, writing a spec, choosing agent or task, timeout, or finding the thread id | [submit.md](references/submit.md) |
| Listing, showing, reading logs, or cancelling tasks | [inspect.md](references/inspect.md) |
| Configuring Fleet or discovering machines | [fleet.md](references/fleet.md) |
| Sending a direct message to a session or thread | [messages.md](references/messages.md) |
| A command exited non-zero, `daemon_unavailable`, or the socket is down | [errors.md](references/errors.md) |
| Supervising shared GPU work or a resource loan | [resource-loans.md](references/resource-loans.md) |
| `homebased` is missing, the daemon is not installed, or the binary was rebuilt | [setup.md](references/setup.md) |

## Minimal flow
Expand Down
27 changes: 23 additions & 4 deletions .agents/skills/homebased/references/errors.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,13 +26,13 @@ Errors go to stderr. With `--json` they are one object:
| `daemon_unavailable` | 1, retryable | Socket missing or refused. | `homebased --json daemon status`. If `socket` is `down`, read setup.md and start or restart the daemon. Tasks already running keep running and still report. |
| `config_invalid` | 2 | An explicitly selected config file is missing, unreadable, or invalid TOML. | Fix the reported path or setting. Run `homebased --json config validate`. |
| `invalid_spec` | 2 | Bad JSON, unknown field, wrong `api_version`, bad UUID, both or neither of `prompt`/`prompt_file`, timeout below 30m, empty command, cross-variant fields, unreadable `prompt_file`, blank prompt text, or a container mount `source` that is missing or exposes a container daemon socket on the machine that runs the task (`input.pointer` is `/workload/mounts/<n>/source`, or `/workload/mounts` when a remote machine refused it). | Fix the field at `input.pointer`. `homebased task schema` prints the schema. |
| `unknown_thread` | 2 | The spec `thread` names no Claude Code session (registry file or transcript) and no Codex thread (session file) on the submitting machine. Checked by `task submit` (also with `--dry-run`), `resource request submit`, and `resource background submit` before any work starts. | Use the exact id from submit.md step 1. `input.suggestions` lists known ids that differ by a likely typo, best first; the message names the best one. A Homebased task id is not a thread. A worker on a remote executor cannot send events to its parent's thread. |
| `unknown_thread` | 2 | The spec `thread` names no Claude Code session (registry file or transcript) and no Codex thread (session file) on the submitting machine. Checked by `task submit`, also with `--dry-run`, before any work starts. | Use the exact id from submit.md step 1. `input.suggestions` lists known ids that differ by a likely typo, best first; the message names the best one. A Homebased task id is not a thread. A worker on a remote executor cannot send events to its parent's thread. |
| `thread_mismatch` | 2 | `HOMEBASED_TASK_ID` is set and the spec `thread` differs from the parent task's thread. `input.parent_task` and `input.parent_thread` name the parent. | Use `input.parent_thread`. Pass `--allow-other-thread` on submit or followup only when the events must go to another session; that thread must still exist. |
| `agent_configuration` | 1 | OpenCode inherited inline configuration is malformed, has the wrong shape, or already defines the generated task agent. | Fix `OPENCODE_CONFIG_CONTENT` without putting credentials in the task spec or logs, then submit again. |
| `invalid_cwd` | 2 | `cwd` is not an existing, accessible host directory on the machine that runs the task. `input.problem` is `not_found`, `not_directory`, or `inaccessible`; `input.value` is the `cwd`. Checked at submit, including `--dry-run`, `resource request submit`, and `resource background submit`; a remote executor or resource authority checks its own file system and refuses at acceptance. | `cwd` is a host path, never a path inside a container. When it names a container mount target, `input.suggested_cwd` is the matching host path under that mount's `source`; use it, and set `workload.workdir` for the directory inside the container. |
| `invalid_cwd` | 2 | `cwd` is not an existing, accessible host directory on the machine that runs the task. `input.problem` is `not_found`, `not_directory`, or `inaccessible`; `input.value` is the `cwd`. Checked at submit, including `--dry-run`; a remote executor checks its own file system and refuses at acceptance. | `cwd` is a host path, never a path inside a container. When it names a container mount target, `input.suggested_cwd` is the matching host path under that mount's `source`; use it, and set `workload.workdir` for the directory inside the container. |
| `executable_missing` | 3 | Requested program missing, not a file, or not executable. Agents also check `HOMEBASED_<AGENT>` overrides. | Submit from a shell where the program is on `PATH`, use an absolute path, or export `HOMEBASED_CODEX`, `HOMEBASED_CLAUDE`, or `HOMEBASED_GROK`. |
| `unknown_dependency` | 3 | An `after` entry names no task submitted through this daemon. `input.task` names it. Checked by `task submit`, also with `--dry-run`. | Use the id of a task submitted from this machine. A task submitted through another machine cannot be a dependency here. |
| `dependency_failed` | 5 | An `after` entry already ended without success, so the task could never start. `input.task` and `input.outcome` name it. `input.outcome` is `failed`, `blocked`, `cancelled`, `lost`, or `unknown`; `unknown` means the task ended but no record says how, as for a task that finished before an update and whose terminal event was pruned. | Handle that task's result first, then submit without it or after its replacement. |
| `dependency_failed` | 5 | An `after` entry already ended without success, so the task could never start. `input.task` and `input.outcome` name it. `input.outcome` is `failed`, `blocked`, `cancelled`, `lost`, or `unknown`; `unknown` means the task ended but no record here says how. | Handle that task's result first, then submit without it or after its replacement. |
| `task_not_found` | 3 | No task with that full UUID exists in the checked scope. | Check the UUID and known Fleet inventory. `task list` shows local tasks only. |
| `cluster_lookup_incomplete` | 1, retryable | A known Fleet machine could not give a definitive task lookup result. | Retry `task show`, `task log`, or `task cancel` after that peer is reachable. `input.unchecked` lists UUIDs that were not checked. This is not proof that the task is absent. |
| `task_unavailable` | 1, retryable | Homebased knows the execution machine, but cannot return its task detail or log. | Restore the executor connection and retry. `task show` can return cached origin state when it has a route. |
Expand All @@ -55,6 +55,7 @@ Errors go to stderr. With `--json` they are one object:
| `duplicate_machine_name` | 5 | More than one live machine has the selected name. | Give each machine a unique `fleet.machine_name`; validate config and rediscover. |
| `cluster_protocol_incompatible` | 5 | The peer and local daemon do not share a supported Fleet protocol version. | Update the incompatible Homebased installation. |
| `submission_outcome_unknown` | 1, retryable | The executor may have accepted the request, but the origin has no definitive reply. | Retry the same spec with the same `--request-id`. Do not create a new request UUID for the same intended task. The error input includes `request_id` and `task_id`. |
| `schema_too_new` | 1 | The database was written by a newer Homebased build. `input.found` and `input.supported` name both schema versions. | Update this installation to the newer build. Do not edit or delete the database. |
| `daemon_busy` | 1 | A daemon call timed out, usually because the disk is slow. The operation may still complete. | Check the current state, for example with `task show` or `task list`, before you repeat a change. A local submit retries this by itself. If it keeps happening, tell the user that the daemon or its disk needs attention. |
| `submission_rejected` | 5 | The executor retained a definitive rejection for this task identity. | Fix the cause and submit again with a new request UUID. |
| `submission_conflict` | 5 | The request UUID was already used with different task content or a different `after` list. | Retry with the original content, or use a new UUID for new work. |
Expand All @@ -68,7 +69,7 @@ This happens after a raw `systemctl stop`, a crash, or an upgrade in progress. W

## Callback failed

For a legacy task row, `callback: "failed"` means its terminal `codex queue` delivery failed; the message text is appended to `<home>/callback-fallback.log`. New sequenced events keep a result for each callback. `task show` lists failed event sequences in `failed_events`; callback delivery failure does not change process status. The origin machine uses the saved callback directory, environment, and resolved Codex path. See [events.md](events.md) for retry limits and origin/executor roles.
`callback: "failed"` means delivery of the terminal event failed; the message text is appended to `<home>/callback-fallback.log`. Every sequenced event keeps its own delivery result. `task show` lists failed event sequences in `failed_events`; callback delivery failure does not change process status. The origin machine uses the saved callback directory, environment, and resolved Codex path. See [events.md](events.md) for retry limits and origin/executor roles.

### Claude Code session events missing

Expand All @@ -85,3 +86,21 @@ Claude Code delivery uses an internal Claude Code socket protocol, not a public
| `message_receiver_unavailable` | 1, retryable | The receiver could not inspect local Codex session metadata. | Check the receiver's session files and retry. |
| `message_invalid` | 2 | A message, UUID, or receiver-side `cwd` failed validation. | Fix the reported value. |
| `message_to_self` | 2 | The selected destination is the source thread, or a `--worker` destination is the source task's own worker. A `--task` destination resolves to its origin thread, not its worker. | Choose a different source or destination. Use `--worker` for a running Claude worker, or `homebased task followup` to resume a finished Codex worker. |

## Queue errors

| Code | Exit | Do this |
| --- | --- | --- |
| `invalid_queue_input` | 2 | Fix the full UUID, level, resource name, step count, or restart window. |
| `move_refused` | 2 | Give one valid placement. Check target existence, state, and level. |
| `job_not_found` | 3 | Check the full job UUID and authority machine. |
| `resource_not_found` | 3 | Read resource list on the target machine. |
| `attention_not_found` | 3 | Refresh resource list. Do not release a different Attention without another machine check. |
| `job_terminal` | 5 | Read the result. Submit a new job only for new work. |
| `job_conflict` | 5 | Restore the original spec for this job UUID. Use a new UUID for a different job. |
| `operation_conflict` | 5 | Restore the original operation content. Use a new UUID for a new action. |
| `resource_conflict` | 5 | Use the existing resource. Names and device indices must be unique. |
| `stale_run` | 5 | Refresh job and resource state. Do not act on an old run. |
| `internal` | 1 | Keep the error and state evidence. Ask the operator to check it. Do not edit the database. |

See [resource-queue.md](resource-queue.md) for job and operation retry identities. A run task in `after` returns `usage` (exit 2). Wait for the job result and submit once its inputs are ready.
7 changes: 5 additions & 2 deletions .agents/skills/homebased/references/events.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,7 @@ The message is one line: the literal prefix `HOMEBASED_EVENT ` followed by one J
| `seq` | Per-task event sequence. Present on new events. Use with `task` as the event identity. It is separate from the `seq` on each worker report. |
| `origin_machine` | Stable UUID of the machine that owns the Codex thread and sends the callback. Present on new events. |
| `execution_machine` | Stable UUID of the machine that ran the child. Present on new events. |
| `name` | Submitted task name. Omitted only for tasks stored before name was required. |
| `display_name` | Non-empty server-derived label: the submitted name, or a workload fallback for unnamed stored rows. |
| `name` | Submitted task name. |
| `workload` | Discriminated union: `{"type":"agent","agent":"…","model":null\|string}`, `{"type":"task","command":[…]}`, or `{"type":"container","image":"…","args":[…]}` with optional `entrypoint` and `gpus`. |
| `thread` | The thread the event was addressed to. |
| `cwd` | The child's working directory. |
Expand All @@ -24,6 +23,10 @@ The message is one line: the literal prefix `HOMEBASED_EVENT ` followed by one J

`process` values: `{"kind":"exit","code":n}`, `{"kind":"signal","signal":n}`, `{"kind":"cancelled"}`, `{"kind":"spawn_failed","message":"…"}`, `{"kind":"runner_lost"}`. Output inactivity never appears as a process result.

## Job events

For `JOB_SUCCEEDED`, `JOB_FAILED`, `JOB_CANCELLED`, `JOB_PREEMPTED`, `JOB_ATTENTION`, `JOB_BLOCKED`, and `JOB_CHECK_DUE`, use [resource-queue.md](resource-queue.md). These events identify a job. Deduplicate by `(job, seq)`. Their flat run fields are explicit nulls when no run exists. Queue runs have no ordinary task callback route. `TASK_PREEMPTED` can appear in task inspection event data for a preempted run; handle the job through `JOB_PREEMPTED`.

## Event table

| `event` | `next_action` | What happened | Do this |
Expand Down
Loading
Loading