Skip to content

Replace GPU loans with a priority queue - #1

Merged
praveenperera merged 31 commits into
masterfrom
gpu-priority-queue
Oct 4, 2026
Merged

praveenperera merged 31 commits into
masterfrom
gpu-priority-queue

Conversation

@praveenperera

@praveenperera praveenperera commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Replaces the GPU resource-loan subsystem with a priority queue that runs any finite command, including Python and other scripts, on an exclusive GPU. Higher-priority work takes the GPU at the running job's next checkpoint, and jobs at the same priority run first in, first out unless someone moves them.

Behavior

  • Resources. The daemon detects GPUs at startup and creates one resource per GPU (gpu0, gpu1, …). A Mac gets gpu0. A machine with no NVIDIA GPU gets an unindexed gpu0 that acts as an exclusive lane. If nvidia-smi is installed but fails, no resource is created on that start, and a later start detects the real devices.
  • One queue per machine. A job targets a machine, local by default. Without a named resource it runs on whichever GPU is free; naming one pins it there.
  • Priorities. high, medium, low, chosen by how urgently the result is needed. Only a strictly higher level preempts. Jobs can be moved within a level and between levels (resource job move).
  • Preemption. Each job must choose preempt:
    • yield: Homebased creates HOMEBASED_YIELD_FILE, and the job saves its state and exits 75.
    • restart: the run is killed and the current step rerun.
    • wait: the job is never interrupted.
    • restart_within (with yield or wait): the run may be restarted while it is younger than the given window.
    • Multi-step jobs get a checkpoint at every step boundary.
  • Checkpoint contract. Each run gets HOMEBASED_JOB_ID, HOMEBASED_JOB_DIR (stable across runs), HOMEBASED_RUN_NUMBER, HOMEBASED_STEP_INDEX, HOMEBASED_RESUME, HOMEBASED_YIELD_FILE, HOMEBASED_RESOURCE, and CUDA_VISIBLE_DEVICES for an indexed GPU. Containers get the same values at /homebased/job and /homebased/run, with Docker selecting the device.
  • Cleanup. After every run, Homebased kills processes the run left behind before the GPU serves again. It finds them by the run's HOMEBASED_TASK_ID in each process's environment, checks each process's identity before every signal, and reports success only after two empty scans a poll apart in which no listed process vanished mid-scan (or eight consecutive empty scans on a busy machine). An unreadable same-user process that started during the run blocks success too. If cleanup cannot finish, the resource is held in Attention until a person checks the machine and runs resource release --attention <id>. Other resources keep serving.
  • Notices and events. The blocked job's thread gets one notice after 15 minutes behind a run asked to yield, or 30 minutes otherwise. The notice never stops anything. Threads receive JOB_SUCCEEDED, JOB_FAILED, JOB_CANCELLED, JOB_PREEMPTED, JOB_BLOCKED, JOB_ATTENTION, and JOB_CHECK_DUE in order and at least once, including across the fleet and after an offline origin returns.
  • Interfaces. homebased resource (list, jobs, register, schema, job submit|show|move|cancel, release), a queue panel and /queue pages in the dashboard, a skill reference (.agents/skills/homebased/references/resource-queue.md), and a runbook (docs/resource-queue.md).
  • Security. Job submit and resource registration are socket-only on the dashboard listener; move, cancel, and release there require same-origin JSON. With fleet enabled, /v1/cluster/* accepts queue requests from peers under the existing trusted LAN and tailnet boundary, the same as task executions; peer authentication remains deferred.

Removed

Loans, supervisors, return work, release watchers, trainer attempts and ownership locks, background launch, and operator and initial-idle attestations, with their CLI, routes, web views, docs, and skill reference. The loan system never ran live work.

Database and fleet protocol

  • The database moves to homebased_v1.sqlite at schema version 1. The old homebased.sqlite is never read, migrated, or deleted. Historical migrations are gone; future migrations go in one place, from version N to N + 1. Opening a database written by a newer build fails with schema_too_new.
  • Stored data no longer repeats facts: a task has one required name, job step and resume state and route cursors are each stored once, priority is stored as a rank, and event and route tables no longer mirror their JSON. The legacy pre-fleet task model is removed.
  • The fleet protocol moves from 2 to 3, so machines on different versions refuse each other instead of misreading data.

Upgrading

  • Update every fleet machine together, because protocol 3 refuses protocol 2.
  • Update each machine when nothing is queued or running. Each starts with an empty task list; tasks the old database tracked send no events, and the old file stays in place.

Refactors

After the reviews below, the branch went through a refactor and simplify pass over the queue code, a per-file refactor of every Rust file, the data-model reshape, a split-and-simplify pass (the task, event, and queue stores, callback delivery, and the large test files split by concern; duplicated code paths merged; unused APIs and restating tests removed), and a pass that fixed bugs those agents reported. Behavior outside the format changes above is unchanged.

Verification

  • just ci passes on Linux and macOS in GitHub Actions and locally: fmt, clippy with -D warnings, web check, lint, tests and build, and all Rust tests.
  • Focused queue tests (actor, routes, fleet, review fixes) each passed three times in a row.
  • Controlled GPU case on the Mac mini with MLX on the Metal GPU: a low-priority yield job ran, saw the yield file after unit 6 and exited 75, a high-priority job ran on gpu0, and the low job resumed at unit 6 as run 2 with HOMEBASED_RESUME=1 and finished all 20 units. JOB_PREEMPTED and JOB_SUCCEEDED callbacks arrived through the real delivery path.
  • Reviews: GPT-6 Astra (high) and Grok 4.7 (high) reviewed the branch. Nine findings were accepted and fixed, each with a test that failed before the fix: a cleanup scan race, the container CUDA ordinal, detection fallback recovery, a natural exit recorded as a cancel, stale preemption decisions, control-directory permissions under a strict umask, stale blocked notices, a direct run-task cancel during a yield, and a misleading listener comment. A fresh Sol review checked the fix pass, and two gaps it found were closed. Greptile's reviews led to more fixes: preemption considers any waiting job that can use a GPU, a reconciled fallback takes its device's name, unreadable processes started during a run hold cleanup, cleanup confirmation tolerates busy machines without missing fork chains. Greptile's score on the reviewed code was 5/5.

Not verified

  • Linux beyond CI. No Linux machine was available; the Linux cleanup code is covered by the CI test suite on ubuntu-latest, not by a GPU host.
  • Real Docker containers and multi-GPU NVIDIA hosts; container behavior is covered by argument and mount contract tests.
  • Fleet submission between physical machines; fleet tests use real daemons on one host.

The loan design never ran live work and is being replaced by a GPU
priority queue that accepts scripts and preempts at checkpoints. This
deletes loans, supervisors, return work, release watchers, trainer
attempts, attestations, their routes, CLI, web views, docs, and skill
references.

Schema version 35 drops the loan tables and the rows those flows wrote
into shared route, cancellation, and message tables. Ordinary tasks are
kept, and every released schema version upgrades to the fresh schema.

Two tests counted callbacks that can arrive at any time after a task
ends; they now count only the callbacks they check.
The GPU priority queue releases a GPU by cleaning up after a run instead
of proving the run stopped. The new cleanup module sweeps every process
for the run's HOMEBASED_TASK_ID, identifies each by pid and start time,
and rereads both before every signal. It also cleans up a lost worker's
process group only when the group can still be attributed to the run.

Run-marker variables are now scrubbed from every process Homebased
spawns, so a marker can never leak into the daemon, a worker, or an
unrelated child. Workers record each child's pid and start time for
lost-worker cleanup, and tasks gain a preempted outcome that never
counts as success.

A CI workflow runs just ci on Linux and macOS, which is where the Linux
cleanup code is tested.
The test cleared the fake agent's sleep control before cancelling, so a
slow agent that had not read the control yet exited 0 first and the
task never reached cancelled. It now clears the control after the
cancel.
Jobs, resources, and runs get typed models: priorities, preemption
modes with restart windows, pinned or any-resource targets, steps, and
each resource's active run lifecycle. A strict job spec validates
structure everywhere and file system checks on the authority, and GPU
detection turns nvidia-smi output or a Mac into resources.

The store keeps one serving order per machine with dense slots, moves,
cancels, operation replay, run reservation, stop causes, cleanup
results, and Attention release, with invariants enforced by
constraints. A run task's terminal commit applies the run classification
in the same transaction and records a job event instead of a task
event. A pure scheduler fills free resources and picks at most one
preemption for the head of the queue.
A queue actor per machine applies the scheduler's decisions: it reserves
and launches runs on free resources with their job directory, control
directory, run variables, CUDA device, and container mounts, commits
stop causes before signalling, and cleans up after every run before the
resource serves again. Cleanup that cannot finish holds the resource in
Attention until a person releases it, while other resources keep
serving.

Blocked and quiet-output notices are recorded as job events in the same
transaction as their state, so a restart neither loses nor repeats one.
The actor rebuilds from stored state after a restart at every point in
the recovery table, and startup now detects GPUs and keeps their
resources.
The message sender harness discarded daemon stderr, so an intermittent
readiness timeout carried no cause. It now keeps stderr and reports
whether the daemon exited.
Jobs can be submitted, moved, cancelled, inspected, and released with
homebased resource commands, from the local machine or across the
fleet. Submission saves an origin job route first, so retries with the
same job ID are safe and a different spec or origin conflicts. Moves,
cancels, and releases carry operation IDs that replay their stored
result.

Job events travel from the authority to the submitting thread in order
and at least once, with cursors committed alongside the events they
track, so an offline origin catches up when it returns. Intermediate
steps send no callback, run tasks send no task callbacks, and an
ordinary task can no longer wait on a run task.
A job's state directory was world-writable for every run. Host steps run
as the directory's owner, so they now get an owner-only directory; only
container steps, whose user may have another UID, get a world-writable
one under the owner-only state root.
Job submit and resource registration were served on the web listener,
where a same-origin check stops browsers but not other clients, so any
host that reached the port could run commands. They are now
socket-only, like task submit; moving, cancelling, and releasing jobs
stay on the web listener behind the same-origin guard.
Agents get a skill reference for choosing priority and preemption, the
checkpoint contract, job specs and commands, job events, Attention, and
cleanup limits, and the skill routes GPU work there. Operators get a
runbook for reading the queue, moving and cancelling jobs, and releasing
a held resource.
free_port releases its port before the daemon binds it, so another
process could take it and the daemon exited with the address in use. A
first start now retries on a fresh port when that happens; restarts
keep their address because peers point at it.
The main page gains a queue panel while any resource is busy or jobs
wait. A queue page shows each resource's state and the machine queue by
level, and a job page shows its steps and runs. Jobs can be moved with
buttons, a keyboard menu, or drag and drop, and cancel and Attention
release ask for confirmation first.
Two reviews of the branch found races and edge cases, each now covered
by a test that failed before its fix:

- cleanup requires a stable empty rescan, so a marked child forked
  during a scan is not missed
- a container sees CUDA ordinal 0 for its one exposed GPU
- a detection fallback is reconciled to its device, and an installed
  but failing nvidia-smi creates no lane until a later start
- a natural exit observed before a stop, or a container that finished
  before the stop was issued, keeps its exit instead of counting as a
  cancel
- a preemption is rechecked and committed in one transaction, so a move
  or cancel that commits first withdraws it
- the run control directory and yield file stay readable to container
  users under a strict umask
- a blocked notice whose episode ended is suppressed at delivery
- cancelling a run task directly records a user cancel before a yield
  exit can requeue the job
@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Too many files!

This PR contains 291 files, which is 191 over the limit of 100.

To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch.

Upgrade to a paid plan to raise the limit.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 5e311d35-d9c4-4d0d-95d7-f8a914bbda9d
📥 Commits

Reviewing files that changed from the base of the PR and between 7593609 and 0c7b003.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (291)
  • .agents/skills/homebased/SKILL.md
  • .agents/skills/homebased/references/errors.md
  • .agents/skills/homebased/references/events.md
  • .agents/skills/homebased/references/inspect.md
  • .agents/skills/homebased/references/messages.md
  • .agents/skills/homebased/references/resource-loans.md
  • .agents/skills/homebased/references/resource-queue.md
  • .agents/skills/homebased/references/setup.md
  • .agents/skills/homebased/references/submit.md
  • .agents/skills/homebased/references/worker.md
  • .github/workflows/ci.yml
  • Cargo.toml
  • README.md
  • docs/resource-loans.md
  • docs/resource-queue.md
  • src/agents.rs
  • src/callback.rs
  • src/callback/claude_inbox.rs
  • src/callback/delivery.rs
  • src/callback/delivery/tests.rs
  • src/callback/destination.rs
  • src/callback/send_check.rs
  • src/callback/tests.rs
  • src/cancellation.rs
  • src/cleanup.rs
  • src/cleanup/linux.rs
  • src/cleanup/macos.rs
  • src/cleanup/tests.rs
  • src/cli.rs
  • src/cli/daemon.rs
  • src/cli/message.rs
  • src/cli/release_watcher.rs
  • src/cli/resource.rs
  • src/cli/task.rs
  • src/client.rs
  • src/config.rs
  • src/container.rs
  • src/container/docker.rs
  • src/container/lifecycle.rs
  • src/container/lifecycle/tests.rs
  • src/container/spec.rs
  • src/curl.rs
  • src/daemon.rs
  • src/daemon/actors.rs
  • src/daemon/actors/callback.rs
  • src/daemon/actors/queue.rs
  • src/daemon/actors/queue/tests.rs
  • src/daemon/actors/resource.rs
  • src/daemon/actors/resource/test_support.rs
  • src/daemon/actors/resource/tests.rs
  • src/daemon/actors/store.rs
  • src/daemon/actors/supervisor.rs
  • src/daemon/actors/supervisor/recovery.rs
  • src/daemon/actors/supervisor/resource_launch.rs
  • src/daemon/actors/supervisor/tests.rs
  • src/daemon/actors/task.rs
  • src/daemon/api.rs
  • src/daemon/api/views.rs
  • src/daemon/cancel_delivery.rs
  • src/daemon/cluster.rs
  • src/daemon/event_sender.rs
  • src/daemon/event_sender/jobs.rs
  • src/daemon/fleet_tasks.rs
  • src/daemon/inspection.rs
  • src/daemon/local_submit.rs
  • src/daemon/message_receiver.rs
  • src/daemon/origin_submit.rs
  • src/daemon/queue_api.rs
  • src/daemon/release_watcher_api.rs
  • src/daemon/resource_action.rs
  • src/daemon/resource_api.rs
  • src/daemon/resource_api/tests.rs
  • src/daemon/resource_background.rs
  • src/daemon/resource_notice_delivery.rs
  • src/daemon/resource_notice_sender.rs
  • src/daemon/resource_submit.rs
  • src/daemon/t3_watch.rs
  • src/daemon/web.rs
  • src/dependency.rs
  • src/digest.rs
  • src/dispatcher_tests.rs
  • src/domain.rs
  • src/error.rs
  • src/events.rs
  • src/files/content.rs
  • src/files/host.rs
  • src/files/listing.rs
  • src/files/token.rs
  • src/fleet/address.rs
  • src/fleet/advertisement.rs
  • src/fleet/directory.rs
  • src/fleet/discovery.rs
  • src/fleet/discovery/mdns.rs
  • src/fleet/discovery/tailscale.rs
  • src/fleet/protocol.rs
  • src/fleet/runtime.rs
  • src/home.rs
  • src/install.rs
  • src/install/launchd.rs
  • src/install/systemd.rs
  • src/invocation.rs
  • src/lib.rs
  • src/machine.rs
  • src/message.rs
  • src/notify.rs
  • src/power/macos.rs
  • src/queue.rs
  • src/queue/checkpoint.rs
  • src/queue/classify.rs
  • src/queue/delivery.rs
  • src/queue/error.rs
  • src/queue/gpu.rs
  • src/queue/schedule.rs
  • src/queue/schedule/tests.rs
  • src/queue/spec.rs
  • src/queue/spec/tests.rs
  • src/queue/tests.rs
  • src/report.rs
  • src/resource.rs
  • src/resource/api.rs
  • src/resource/background_launch.rs
  • src/resource/bound_action.rs
  • src/resource/command_shape.rs
  • src/resource/foreground.rs
  • src/resource/id.rs
  • src/resource/initial_idle.rs
  • src/resource/operator_release.rs
  • src/resource/ownership_lock.rs
  • src/resource/ownership_lock/attempt_evidence.rs
  • src/resource/ownership_lock/lock_probe.rs
  • src/resource/release_checkpoint.rs
  • src/resource/release_watcher.rs
  • src/resource/return_window.rs
  • src/resource/store.rs
  • src/resource/store/acceptance.rs
  • src/resource/store/assigned_task.rs
  • src/resource/store/cancellation.rs
  • src/resource/store/checkpoint.rs
  • src/resource/store/codec.rs
  • src/resource/store/error.rs
  • src/resource/store/notice.rs
  • src/resource/store/provenance.rs
  • src/resource/store/queue.rs
  • src/resource/store/release_completion.rs
  • src/resource/store/release_loan.rs
  • src/resource/store/release_watcher.rs
  • src/resource/store/resources.rs
  • src/resource/store/return_window.rs
  • src/resource/store/revision.rs
  • src/resource/store/rows.rs
  • src/resource/store/schema.rs
  • src/resource/store/test_support.rs
  • src/resource/store/tests.rs
  • src/resource/trainer_publication.rs
  • src/run_env.rs
  • src/runner.rs
  • src/runner/container.rs
  • src/spec.rs
  • src/spec/host.rs
  • src/spec/tests.rs
  • src/store.rs
  • src/store/admission.rs
  • src/store/cancellation.rs
  • src/store/container.rs
  • src/store/dependency.rs
  • src/store/events.rs
  • src/store/events/inbox.rs
  • src/store/events/outbox.rs
  • src/store/events/retention.rs
  • src/store/events/tests.rs
  • src/store/identity.rs
  • src/store/identity/tests.rs
  • src/store/lifecycle.rs
  • src/store/message.rs
  • src/store/presentation.rs
  • src/store/queue.rs
  • src/store/queue/delivery.rs
  • src/store/queue/delivery/tests.rs
  • src/store/queue/interface.rs
  • src/store/queue/jobs.rs
  • src/store/queue/resources.rs
  • src/store/queue/rows.rs
  • src/store/queue/runs.rs
  • src/store/queue/runtime.rs
  • src/store/queue/tests.rs
  • src/store/resource.rs
  • src/store/resource/action_task.rs
  • src/store/resource/assigned_task.rs
  • src/store/resource/background.rs
  • src/store/resource/cancellation.rs
  • src/store/resource/controls.rs
  • src/store/resource/controls/test_support.rs
  • src/store/resource/initial_idle.rs
  • src/store/resource/operator_release.rs
  • src/store/resource/release_checkpoint.rs
  • src/store/resource/release_completion.rs
  • src/store/resource/release_proof.rs
  • src/store/resource/release_watcher.rs
  • src/store/resource/restore.rs
  • src/store/resource/test_support.rs
  • src/store/resource/tests.rs
  • src/store/resource/tests/assigned_task.rs
  • src/store/resource/tests/background.rs
  • src/store/resource/tests/cancellation.rs
  • src/store/resource/tests/completed_boundary.rs
  • src/store/resource/tests/container.rs
  • src/store/resource/tests/controls.rs
  • src/store/resource/tests/docker.rs
  • src/store/resource/tests/ended_trainer.rs
  • src/store/resource/tests/fixtures.rs
  • src/store/resource/tests/foreground.rs
  • src/store/resource/tests/initial_idle.rs
  • src/store/resource/tests/operator_release.rs
  • src/store/resource/tests/queue_authority.rs
  • src/store/resource/tests/registration.rs
  • src/store/resource/tests/release_checkpoint.rs
  • src/store/resource/tests/release_completion.rs
  • src/store/resource/tests/release_transaction.rs
  • src/store/resource/tests/release_watcher.rs
  • src/store/resource/tests/release_watcher_acceptance.rs
  • src/store/resource/tests/release_watcher_binding.rs
  • src/store/resource/tests/restore.rs
  • src/store/resource/tests/return_deadline.rs
  • src/store/resource/tests/supervised_launch.rs
  • src/store/resource/tests/trainer_association.rs
  • src/store/resource/trainer_association.rs
  • src/store/resource/trainer_lock.rs
  • src/store/schema.rs
  • src/store/task.rs
  • src/store/tests.rs
  • src/submission.rs
  • src/t3.rs
  • src/t3/rpc.rs
  • src/t3/v1.rs
  • src/t3/v2.rs
  • src/thread_title.rs
  • src/update.rs
  • tests/cli_smoke.rs
  • tests/fleet.rs
  • tests/fleet/cancellation.rs
  • tests/fleet/dependencies.rs
  • tests/fleet/events.rs
  • tests/fleet/inspection.rs
  • tests/fleet/messages.rs
  • tests/fleet/peers.rs
  • tests/fleet/resource_jobs.rs
  • tests/fleet/submission.rs
  • tests/integration.rs
  • tests/integration/containers.rs
  • tests/integration/dashboard.rs
  • tests/integration/dependencies.rs
  • tests/integration/lifecycle.rs
  • tests/integration/reports.rs
  • tests/integration/resource_jobs.rs
  • tests/integration/service.rs
  • tests/integration/source_rules.rs
  • tests/integration/submission.rs
  • tests/integration/workers.rs
  • tests/message_receiver.rs
  • tests/message_sender.rs
  • web/src/lib/api.ts
  • web/src/lib/components/CallbackBadge.svelte
  • web/src/lib/components/ConfirmDialog.svelte
  • web/src/lib/components/ExpandToggle.svelte
  • web/src/lib/components/JobControls.svelte
  • web/src/lib/components/LevelBadge.svelte
  • web/src/lib/components/QueueBadge.svelte
  • web/src/lib/components/QueuePanel.svelte
  • web/src/lib/components/ReleaseButton.svelte
  • web/src/lib/components/ResourceQueuePanel.svelte
  • web/src/lib/components/StatusBadge.svelte
  • web/src/lib/components/TaskList.svelte
  • web/src/lib/daemon.svelte.ts
  • web/src/lib/format.ts
  • web/src/lib/queue-view.test.ts
  • web/src/lib/queue-view.ts
  • web/src/lib/resource-actions.ts
  • web/src/lib/resource-operations.svelte.ts
  • web/src/lib/resource-operations.test.ts
  • web/src/lib/resource-operations.ts
  • web/src/lib/resource-state.test.ts
  • web/src/lib/resource-state.ts
  • web/src/lib/resources.ts
  • web/src/routes/+page.svelte
  • web/src/routes/files/+page.svelte
  • web/src/routes/queue/+page.svelte
  • web/src/routes/queue/jobs/[id]/+page.svelte
  • web/src/routes/resources/+page.svelte
  • web/src/routes/resources/[id]/+page.svelte
  • web/src/routes/tasks/[id]/+page.svelte
  • xtask/src/bump.rs

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • Review on demand using usage pricing
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@greptile-apps

greptile-apps Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Replaces GPU resource system with priority queue implementation.

The PR appears safe to merge based on the reviewed changes and resolved findings.

Summary

The PR replaces GPU loans with a priority-based job queue, checkpoint-aware preemption, cleanup gating, fleet support, and queue interfaces. Since the previous review, it also makes T3 delivery intent writes atomic and adds failure-path tests. All previous Greptile threads are resolved; no new actionable issue was established.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Submit[Submit job] --> Queue[Machine priority queue]
  Queue --> GPU[Exclusive GPU resource]
  GPU --> Run[Run step]
  Run --> Cleanup[Clean up run processes]
  Cleanup -->|Confirmed| Queue
  Cleanup -->|Incomplete| Attention[Hold resource in Attention]
Loading

Reviews (5) · Last reviewed commit: "Expect quoted agent paths in the systemd..."

Comment thread src/cleanup.rs
Comment thread src/queue/schedule.rs Outdated
Comment thread src/store/queue.rs Outdated
Comment thread src/store.rs Outdated
The macOS reader reports an unreaped zombie as exited with the identity
it ran under, but the Linux reader returned it as a found process
marked zombie, which left the exited case unused on Linux and failed
clippy there. Linux now reports zombies the same way; cleanup already
treated both forms alike.
The test assumed macOS hides an Apple binary's environment, which holds
with System Integrity Protection on but not on CI runners without it,
where the sweep correctly finds and stops the process. The test now
checks the right outcome for whichever the host reports.
The cleanup escape test waited for the identity of a helper that exits
as soon as it records its orphan, which a fast Linux runner reaps into
a zombie first; it now waits only for the orphan. Fleet test daemons
default to a no-op codex, because submit resolves codex for callbacks
and CI hosts do not have it.
Preemption considered only the head of the queue, so an urgent job that
could run on any GPU waited behind a head pinned to a busy one. The
first queued job with an eligible lower-priority run now gets the
preemption, as launches already skip a pinned job that cannot start.
A reconciled detection fallback also takes the name of its device, so
a job that names that GPU is accepted.
The upgrade tests rebuild old schemas from the current one. A database
written by released v0.13, whose schema v0.14 shares, now upgrades in a
test as well, keeping its finished task and events.
Cleanup treated any process that exited mid-scan as an unstable scan,
so on a machine where short-lived processes exit constantly it could
never confirm and would hold the GPU in Attention. Two empty scans a
poll apart are enough to catch a child left by a parent that forked and
exited during a scan, so exits no longer reset the confirmation.
Cleanup ignored any same-user process whose environment it could not
read, so a run's detached child that made itself unreadable could keep
using the GPU after the resource was released. A leftover must have
started after the run's workload, so an unreadable same-user process
that started then now blocks confirmation, and if it remains the
resource goes to Attention naming it. Older processes, such as agents
that refuse environment reads, stay ignored.
Each release that introduced a schema version, v0.4.0 through v0.13.0,
now has a fixture dumped from a database its own binary wrote, with one
finished task. A test upgrades each to the fresh schema and keeps the
task, instead of relying only on schemas rebuilt from the current one.
@praveenperera

Copy link
Copy Markdown
Member Author

@greptileai review

Comment thread src/cleanup.rs Outdated
A marked process that forks and exits while a scan inspects it can leave
its child out of that scan, and a second generation could do the same in
the next scan, so two empty scans could release the GPU too early. A
scan in which a listed process exited before inspection no longer
counts toward the two quiet empty scans. A run of eight empty scans
still finishes cleanup on a busy machine where short-lived processes
exit during every scan.
@praveenperera

Copy link
Copy Markdown
Member Author

@greptileai review

The store writes job state from typed JobState values, steps are matched
as StepWorkload without impossible agent arms, cleanup has a single
zombie form, queue error codes come from one mapping, the checkpoint
contract names its host and container forms, and the job spec reuses
the task spec's helpers. Tests drop the phase and review prefixes from
how the queue was built, and the phase5 test files become
resource_jobs. Behavior, formats, and error codes are unchanged.

Queue integration tests also wait up to 30 seconds for a step or job to
start, since cleanup takes several scans while parallel tests keep
starting and exiting processes.
Dead helpers and variants are gone, branch code that was duplicated now
shares helpers, and tests that only restated the implementation (level
order and error-code tables, redundant classification and detection
cases) are removed or merged into table tests. Test names no longer
refer to build phases or review rounds.
A per-file refactor pass flattened nested control flow, extracted
private helpers for repeated code, gave unnamed tuples named types,
replaced wildcard imports, and fixed misplaced or redundant comments.
Each change stays inside one file and keeps behavior and interfaces.
The database moves to homebased_v1.sqlite at schema version 1 with one
place for future migrations; the old homebased.sqlite is never read, and
a database from a newer build is refused. Historical migrations, the
pre-fleet task model, and release fixtures are gone.

Stored and wire data no longer repeat facts: a task has one required
name, job step and resume state and route cursors are each stored once,
priority is stored as a rank, and event and route tables stop mirroring
their JSON. CLI queue output is typed. The fleet protocol moves to 3, so
mixed versions refuse each other.

A finished queue run no longer reads as a pending callback, which made
daemon stop wait until it timed out. A transport retry test now makes
its accepted socket blocking, since macOS streams inherit the
listener's nonblocking mode.
The task store splits into task, admission, lifecycle, and presentation
modules, events into retention, outbox, and inbox, and the queue store
into resources, jobs, runs, and rows; spec host checks, callback
delivery, and the large integration and fleet test files split by
theme. Twin callback send paths, layered exit and report wrappers,
duplicate identity helpers, and dashboard types that repeated their
schemas collapse into one each, and unused APIs and restating tests are
gone.
Agents refactoring each file reported 24 bugs; each is now fixed with a
test that failed before the fix. Among them: the macOS process argument
parser could skip an environment entry and hide a run marker from
cleanup, a panicked delivery worker kept its task or job key active
until restart, daemon stop could report success while the daemon ran,
out-of-range ports and malformed Host values were accepted, a null
CFString could be released, log tails read whole files, systemd unit
paths were unquoted, and several tests could pass without checking
their behavior.
@praveenperera

Copy link
Copy Markdown
Member Author

@greptileai review

The intent record was written straight to its final path, so an
interrupted or failed write left a record that every retry failed on,
blocking a job's callbacks for good, and a failed sync could be skipped
by the next attempt. It is now written to a temporary file, synced,
renamed, and its directory synced before dispatch; retries remove
leftover temporary files and redo the sync for an existing record.
The unit now quotes agent paths so a path with spaces stays one value,
and this Linux-only test still expected the unquoted form.
@praveenperera

Copy link
Copy Markdown
Member Author

@greptileai review

@praveenperera
praveenperera merged commit a65d660 into master Oct 4, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant