Replace GPU loans with a priority queue - #1
Conversation
The loan design never ran live work and is being replaced by a GPU priority queue that accepts scripts and preempts at checkpoints. This deletes loans, supervisors, return work, release watchers, trainer attempts, attestations, their routes, CLI, web views, docs, and skill references. Schema version 35 drops the loan tables and the rows those flows wrote into shared route, cancellation, and message tables. Ordinary tasks are kept, and every released schema version upgrades to the fresh schema. Two tests counted callbacks that can arrive at any time after a task ends; they now count only the callbacks they check.
The GPU priority queue releases a GPU by cleaning up after a run instead of proving the run stopped. The new cleanup module sweeps every process for the run's HOMEBASED_TASK_ID, identifies each by pid and start time, and rereads both before every signal. It also cleans up a lost worker's process group only when the group can still be attributed to the run. Run-marker variables are now scrubbed from every process Homebased spawns, so a marker can never leak into the daemon, a worker, or an unrelated child. Workers record each child's pid and start time for lost-worker cleanup, and tasks gain a preempted outcome that never counts as success. A CI workflow runs just ci on Linux and macOS, which is where the Linux cleanup code is tested.
The test cleared the fake agent's sleep control before cancelling, so a slow agent that had not read the control yet exited 0 first and the task never reached cancelled. It now clears the control after the cancel.
Jobs, resources, and runs get typed models: priorities, preemption modes with restart windows, pinned or any-resource targets, steps, and each resource's active run lifecycle. A strict job spec validates structure everywhere and file system checks on the authority, and GPU detection turns nvidia-smi output or a Mac into resources. The store keeps one serving order per machine with dense slots, moves, cancels, operation replay, run reservation, stop causes, cleanup results, and Attention release, with invariants enforced by constraints. A run task's terminal commit applies the run classification in the same transaction and records a job event instead of a task event. A pure scheduler fills free resources and picks at most one preemption for the head of the queue.
A queue actor per machine applies the scheduler's decisions: it reserves and launches runs on free resources with their job directory, control directory, run variables, CUDA device, and container mounts, commits stop causes before signalling, and cleans up after every run before the resource serves again. Cleanup that cannot finish holds the resource in Attention until a person releases it, while other resources keep serving. Blocked and quiet-output notices are recorded as job events in the same transaction as their state, so a restart neither loses nor repeats one. The actor rebuilds from stored state after a restart at every point in the recovery table, and startup now detects GPUs and keeps their resources.
The message sender harness discarded daemon stderr, so an intermittent readiness timeout carried no cause. It now keeps stderr and reports whether the daemon exited.
Jobs can be submitted, moved, cancelled, inspected, and released with homebased resource commands, from the local machine or across the fleet. Submission saves an origin job route first, so retries with the same job ID are safe and a different spec or origin conflicts. Moves, cancels, and releases carry operation IDs that replay their stored result. Job events travel from the authority to the submitting thread in order and at least once, with cursors committed alongside the events they track, so an offline origin catches up when it returns. Intermediate steps send no callback, run tasks send no task callbacks, and an ordinary task can no longer wait on a run task.
A job's state directory was world-writable for every run. Host steps run as the directory's owner, so they now get an owner-only directory; only container steps, whose user may have another UID, get a world-writable one under the owner-only state root.
Job submit and resource registration were served on the web listener, where a same-origin check stops browsers but not other clients, so any host that reached the port could run commands. They are now socket-only, like task submit; moving, cancelling, and releasing jobs stay on the web listener behind the same-origin guard.
Agents get a skill reference for choosing priority and preemption, the checkpoint contract, job specs and commands, job events, Attention, and cleanup limits, and the skill routes GPU work there. Operators get a runbook for reading the queue, moving and cancelling jobs, and releasing a held resource.
free_port releases its port before the daemon binds it, so another process could take it and the daemon exited with the address in use. A first start now retries on a fresh port when that happens; restarts keep their address because peers point at it.
The main page gains a queue panel while any resource is busy or jobs wait. A queue page shows each resource's state and the machine queue by level, and a job page shows its steps and runs. Jobs can be moved with buttons, a keyboard menu, or drag and drop, and cancel and Attention release ask for confirmation first.
Two reviews of the branch found races and edge cases, each now covered by a test that failed before its fix: - cleanup requires a stable empty rescan, so a marked child forked during a scan is not missed - a container sees CUDA ordinal 0 for its one exposed GPU - a detection fallback is reconciled to its device, and an installed but failing nvidia-smi creates no lane until a later start - a natural exit observed before a stop, or a container that finished before the stop was issued, keeps its exit instead of counting as a cancel - a preemption is rechecked and committed in one transaction, so a move or cancel that commits first withdraws it - the run control directory and yield file stay readable to container users under a strict umask - a blocked notice whose episode ended is suppressed at delivery - cancelling a run task directly records a user cancel before a yield exit can requeue the job
|
Important Review skippedToo many files! This PR contains 291 files, which is 191 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. ⚙️ Run configuration
⛔ Files ignored due to path filters (1)
📒 Files selected for processing (291)
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
The macOS reader reports an unreaped zombie as exited with the identity it ran under, but the Linux reader returned it as a found process marked zombie, which left the exited case unused on Linux and failed clippy there. Linux now reports zombies the same way; cleanup already treated both forms alike.
The test assumed macOS hides an Apple binary's environment, which holds with System Integrity Protection on but not on CI runners without it, where the sweep correctly finds and stops the process. The test now checks the right outcome for whichever the host reports.
The cleanup escape test waited for the identity of a helper that exits as soon as it records its orphan, which a fast Linux runner reaps into a zombie first; it now waits only for the orphan. Fleet test daemons default to a no-op codex, because submit resolves codex for callbacks and CI hosts do not have it.
Preemption considered only the head of the queue, so an urgent job that could run on any GPU waited behind a head pinned to a busy one. The first queued job with an eligible lower-priority run now gets the preemption, as launches already skip a pinned job that cannot start. A reconciled detection fallback also takes the name of its device, so a job that names that GPU is accepted.
The upgrade tests rebuild old schemas from the current one. A database written by released v0.13, whose schema v0.14 shares, now upgrades in a test as well, keeping its finished task and events.
Cleanup treated any process that exited mid-scan as an unstable scan, so on a machine where short-lived processes exit constantly it could never confirm and would hold the GPU in Attention. Two empty scans a poll apart are enough to catch a child left by a parent that forked and exited during a scan, so exits no longer reset the confirmation.
Cleanup ignored any same-user process whose environment it could not read, so a run's detached child that made itself unreadable could keep using the GPU after the resource was released. A leftover must have started after the run's workload, so an unreadable same-user process that started then now blocks confirmation, and if it remains the resource goes to Attention naming it. Older processes, such as agents that refuse environment reads, stay ignored.
Each release that introduced a schema version, v0.4.0 through v0.13.0, now has a fixture dumped from a database its own binary wrote, with one finished task. A test upgrades each to the fresh schema and keeps the task, instead of relying only on schemas rebuilt from the current one.
|
@greptileai review |
A marked process that forks and exits while a scan inspects it can leave its child out of that scan, and a second generation could do the same in the next scan, so two empty scans could release the GPU too early. A scan in which a listed process exited before inspection no longer counts toward the two quiet empty scans. A run of eight empty scans still finishes cleanup on a busy machine where short-lived processes exit during every scan.
|
@greptileai review |
The store writes job state from typed JobState values, steps are matched as StepWorkload without impossible agent arms, cleanup has a single zombie form, queue error codes come from one mapping, the checkpoint contract names its host and container forms, and the job spec reuses the task spec's helpers. Tests drop the phase and review prefixes from how the queue was built, and the phase5 test files become resource_jobs. Behavior, formats, and error codes are unchanged. Queue integration tests also wait up to 30 seconds for a step or job to start, since cleanup takes several scans while parallel tests keep starting and exiting processes.
Dead helpers and variants are gone, branch code that was duplicated now shares helpers, and tests that only restated the implementation (level order and error-code tables, redundant classification and detection cases) are removed or merged into table tests. Test names no longer refer to build phases or review rounds.
A per-file refactor pass flattened nested control flow, extracted private helpers for repeated code, gave unnamed tuples named types, replaced wildcard imports, and fixed misplaced or redundant comments. Each change stays inside one file and keeps behavior and interfaces.
The database moves to homebased_v1.sqlite at schema version 1 with one place for future migrations; the old homebased.sqlite is never read, and a database from a newer build is refused. Historical migrations, the pre-fleet task model, and release fixtures are gone. Stored and wire data no longer repeat facts: a task has one required name, job step and resume state and route cursors are each stored once, priority is stored as a rank, and event and route tables stop mirroring their JSON. CLI queue output is typed. The fleet protocol moves to 3, so mixed versions refuse each other. A finished queue run no longer reads as a pending callback, which made daemon stop wait until it timed out. A transport retry test now makes its accepted socket blocking, since macOS streams inherit the listener's nonblocking mode.
The task store splits into task, admission, lifecycle, and presentation modules, events into retention, outbox, and inbox, and the queue store into resources, jobs, runs, and rows; spec host checks, callback delivery, and the large integration and fleet test files split by theme. Twin callback send paths, layered exit and report wrappers, duplicate identity helpers, and dashboard types that repeated their schemas collapse into one each, and unused APIs and restating tests are gone.
Agents refactoring each file reported 24 bugs; each is now fixed with a test that failed before the fix. Among them: the macOS process argument parser could skip an environment entry and hide a run marker from cleanup, a panicked delivery worker kept its task or job key active until restart, daemon stop could report success while the daemon ran, out-of-range ports and malformed Host values were accepted, a null CFString could be released, log tails read whole files, systemd unit paths were unquoted, and several tests could pass without checking their behavior.
|
@greptileai review |
The intent record was written straight to its final path, so an interrupted or failed write left a record that every retry failed on, blocking a job's callbacks for good, and a failed sync could be skipped by the next attempt. It is now written to a temporary file, synced, renamed, and its directory synced before dispatch; retries remove leftover temporary files and redo the sync for an existing record.
The unit now quotes agent paths so a path with spaces stays one value, and this Linux-only test still expected the unquoted form.
|
@greptileai review |
Replaces the GPU resource-loan subsystem with a priority queue that runs any finite command, including Python and other scripts, on an exclusive GPU. Higher-priority work takes the GPU at the running job's next checkpoint, and jobs at the same priority run first in, first out unless someone moves them.
Behavior
gpu0,gpu1, …). A Mac getsgpu0. A machine with no NVIDIA GPU gets an unindexedgpu0that acts as an exclusive lane. Ifnvidia-smiis installed but fails, no resource is created on that start, and a later start detects the real devices.high,medium,low, chosen by how urgently the result is needed. Only a strictly higher level preempts. Jobs can be moved within a level and between levels (resource job move).preempt:yield: Homebased createsHOMEBASED_YIELD_FILE, and the job saves its state and exits 75.restart: the run is killed and the current step rerun.wait: the job is never interrupted.restart_within(withyieldorwait): the run may be restarted while it is younger than the given window.HOMEBASED_JOB_ID,HOMEBASED_JOB_DIR(stable across runs),HOMEBASED_RUN_NUMBER,HOMEBASED_STEP_INDEX,HOMEBASED_RESUME,HOMEBASED_YIELD_FILE,HOMEBASED_RESOURCE, andCUDA_VISIBLE_DEVICESfor an indexed GPU. Containers get the same values at/homebased/joband/homebased/run, with Docker selecting the device.HOMEBASED_TASK_IDin each process's environment, checks each process's identity before every signal, and reports success only after two empty scans a poll apart in which no listed process vanished mid-scan (or eight consecutive empty scans on a busy machine). An unreadable same-user process that started during the run blocks success too. If cleanup cannot finish, the resource is held inAttentionuntil a person checks the machine and runsresource release --attention <id>. Other resources keep serving.JOB_SUCCEEDED,JOB_FAILED,JOB_CANCELLED,JOB_PREEMPTED,JOB_BLOCKED,JOB_ATTENTION, andJOB_CHECK_DUEin order and at least once, including across the fleet and after an offline origin returns.homebased resource(list,jobs,register,schema,job submit|show|move|cancel,release), a queue panel and/queuepages in the dashboard, a skill reference (.agents/skills/homebased/references/resource-queue.md), and a runbook (docs/resource-queue.md)./v1/cluster/*accepts queue requests from peers under the existing trusted LAN and tailnet boundary, the same as task executions; peer authentication remains deferred.Removed
Loans, supervisors, return work, release watchers, trainer attempts and ownership locks, background launch, and operator and initial-idle attestations, with their CLI, routes, web views, docs, and skill reference. The loan system never ran live work.
Database and fleet protocol
homebased_v1.sqliteat schema version 1. The oldhomebased.sqliteis never read, migrated, or deleted. Historical migrations are gone; future migrations go in one place, from version N to N + 1. Opening a database written by a newer build fails withschema_too_new.name, job step and resume state and route cursors are each stored once, priority is stored as a rank, and event and route tables no longer mirror their JSON. The legacy pre-fleet task model is removed.Upgrading
Refactors
After the reviews below, the branch went through a refactor and simplify pass over the queue code, a per-file refactor of every Rust file, the data-model reshape, a split-and-simplify pass (the task, event, and queue stores, callback delivery, and the large test files split by concern; duplicated code paths merged; unused APIs and restating tests removed), and a pass that fixed bugs those agents reported. Behavior outside the format changes above is unchanged.
Verification
just cipasses on Linux and macOS in GitHub Actions and locally: fmt, clippy with-D warnings, web check, lint, tests and build, and all Rust tests.yieldjob ran, saw the yield file after unit 6 and exited 75, a high-priority job ran ongpu0, and the low job resumed at unit 6 as run 2 withHOMEBASED_RESUME=1and finished all 20 units.JOB_PREEMPTEDandJOB_SUCCEEDEDcallbacks arrived through the real delivery path.Not verified
ubuntu-latest, not by a GPU host.