Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
42 commits
Select commit Hold shift + click to select a range
9331ef6
feat: add the execution backend contract
florian-simvia Sep 14, 2026
a60d8d7
feat: add a scripted fake execution backend
florian-simvia Sep 14, 2026
85cbcdc
feat: launch a case through an execution backend
florian-simvia Sep 14, 2026
19a168e
fix: do not finalize a case an execution backend owns
florian-simvia Sep 14, 2026
aa71896
feat: add the backend synchronisation pass
florian-simvia Sep 14, 2026
fa177e8
feat: poll execution backends from a background loop
florian-simvia Sep 14, 2026
09aa474
feat: cancel and synchronise backend cases from the CLI and the kill …
florian-simvia Sep 14, 2026
055f059
test: cover the remote execution lifecycle end to end
florian-simvia Sep 14, 2026
324aabc
fix: do not report the reserved registry marker as a case
florian-simvia Sep 14, 2026
ed53536
feat: let an adapter declare its observability files
florian-simvia Sep 14, 2026
e66bd48
feat: add the pure Qarnot helpers
florian-simvia Sep 14, 2026
590ba08
feat: add the [qarnot] settings table and the optional extra
florian-simvia Sep 14, 2026
5604d7a
feat: submit a case to Qarnot through the docker-batch profile
florian-simvia Sep 14, 2026
ee54c8a
feat: observe, download and cancel a Qarnot task
florian-simvia Sep 14, 2026
f8f11a0
feat: report the Qarnot prerequisites in doctor, and document the bac…
florian-simvia Sep 14, 2026
ca949b3
feat: list the execution backends in /api/app_config
florian-simvia Sep 15, 2026
5a88864
feat: choose the execution backend per launch over HTTP
florian-simvia Sep 15, 2026
cf2a696
feat: choose where a launch runs, from the Run dialog
florian-simvia Sep 15, 2026
424a4df
feat: show what a cloud case used, in the status table
florian-simvia Sep 15, 2026
c424b0a
feat: rebuild the dashboard and document the cloud launch
florian-simvia Sep 15, 2026
4efe44a
test: add the live Qarnot round trip, skipped by default
florian-simvia Sep 15, 2026
aa2fc13
fix: do not fail when stopping a backend task that already finished
florian-simvia Sep 15, 2026
b2ba542
feat: report the Qarnot account quotas in doctor
florian-simvia Sep 15, 2026
ecb18c2
fix: do not hold the registry lock across a backend submit
florian-simvia Sep 15, 2026
c0682e2
fix: make the registry lock re-entrant within a thread
florian-simvia Sep 15, 2026
fefb176
fix: ask Qarnot for enough cores to run the requested ranks
florian-simvia Sep 15, 2026
7af5654
feat: let an execution backend declare its launch options
florian-simvia Sep 15, 2026
c07dc62
feat: carry per-launch options through to the backend
florian-simvia Sep 15, 2026
d1f3374
feat: add the Qarnot scheduling and node constraint values
florian-simvia Sep 15, 2026
4853fee
feat: honour the chosen scheduling and node when submitting
florian-simvia Sep 15, 2026
fb4dc9e
feat: list the account node types, cached and degradable
florian-simvia Sep 15, 2026
f74eab4
feat: serve a backend launch option catalogue to the dialog
florian-simvia Sep 15, 2026
61e6036
feat: choose a backend launch options from the Run dialog
florian-simvia Sep 15, 2026
27e89f3
fix: only send hardware constraints the account offers
florian-simvia Sep 15, 2026
fd6cb4d
fix: lay the remote case out as a study, so the solver finds its mesh
florian-simvia Sep 15, 2026
cc0ce76
fix: let the adapter tell its solver where the shared dirs are
florian-simvia Sep 15, 2026
cd15e1b
fix: plot live residuals against the real time step
florian-simvia Sep 15, 2026
d7e6e7e
feat: make Refresh pull fresh cloud data, and unfreeze the RESU size
florian-simvia Sep 15, 2026
4188a61
fix: keep the plot panels usable when a case has no results
florian-simvia Sep 15, 2026
f561cd4
refactor: drop the Cloud Time and Cloud Cores columns
florian-simvia Sep 15, 2026
fb8cc20
fix: report the iteration a finished case actually reached
florian-simvia Sep 15, 2026
400fbc6
fix: trust the failure marker over the log heuristic
florian-simvia Sep 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,11 +8,30 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/), and this

Make the dashboard follow what each solver can actually do: panels and action buttons are now derived from the solver adapter instead of being shown for every solver.

### Added
- `POST /api/sync_backends`, and a Refresh button that uses it: for a case an execution backend owns, Refresh now asks the server to pull fresh data before rereading. It used to reread local files that only change when the background sync pass runs, so the panels looked frozen between two passes. Throttled to one pass every ten seconds and never failing, so clicking repeatedly costs nothing and a provider outage does not break the button. `csauto status` shares the same throttle
- `ExecutionBackend` boundary (`csauto/backends/`): a case can be launched on a remote execution service instead of a local process or a Slurm job. A backend implements five verbs (`submit`, `poll`, `sync`, `fetch_final`, `cancel`) and translates its own state names into csauto statuses; the core never sees a provider's API. The shipped `fake` backend covers the whole lifecycle in tests with no network, the way the `stub` solver does without a solver. Two new settings, `backend_poll_interval_s` and `backend_sync_interval_s`
- Qarnot cloud execution (`backend = "qarnot"`, optional `[qarnot]` extra): a campaign can run on your own Qarnot account through the generic `docker-batch` profile, with the same docker image used locally. The shared mesh directories travel once per campaign and the case inputs once per case; while a run is in progress only the files the solver adapter declares in `observability_globs` come back, so residuals, probes and logs are live in the dashboard without downloading a multi-gigabyte results directory. The token is read from `QARNOT_TOKEN` and is rejected if found in `csauto.toml`. `csauto doctor --backend qarnot` reports each prerequisite. code_saturne and the stub solver only; code_aster builds no remote command yet
- A **Run on** selector in the dashboard's Run dialog: a launch can be sent to an execution backend instead of the machine hosting the server, chosen per launch rather than per campaign, so one case can be verified locally while the rest go to the cloud. Choosing a backend restates the case count and that the run is billed to your own account before the button is pressed. The list comes from `GET /api/app_config`, so the frontend names no provider; `POST /api/run_case` accepts a matching `backend` field.
- Per-launch execution options: a backend declares what a user may choose and the Run dialog renders it, so choosing Qarnot now offers a scheduling **Priority** (`Flex`, `OnDemand`, `Reserved`) and a **Node type** read from your own account. The core carries the chosen values without interpreting them, which keeps a provider's vocabulary inside its own module, and records them in the case history. The catalogue comes from a dedicated `GET /api/launch_options`, called when the dialog opens rather than on every page load, cached for ten minutes, and degraded rather than failed when the provider is unreachable: a launch never depends on it

### Changed
- Dashboard panels and action buttons are derived from what the solver adapter implements, rather than declared: a panel appears when the adapter provides what feeds it (`find_residuals_files`, `list_probe_files`, or a non-empty `compare_kinds` / `performance_columns` / `control_actions`), and the Restart, Stop, control and Open GUI controls follow the same rule. Solvers other than code_saturne lose the panels and buttons they could never feed: code_aster and the stub solver now show Status, Compare, Log Tail and Recent Errors only. code_saturne is unchanged. Adapters can no longer declare `dashboard_panels`, which now raises `TypeError` at import time
- `csauto doctor` reports the panels and capabilities derived for the configured solver

### Fixed
- A run killed by code_saturne's runaway-computation check was reported `DONE`. The solver prints its closing banner before aborting and writes the real message to `error`, not to the log, so the log heuristic read the run as a success; the explicit `run_status.failed` marker beside it was only consulted when the log gave no verdict at all. The marker now decides, and the log is the fallback
- LAST ITER stopped at the iteration a run was at partway through. code_saturne writes a `run_status.running` marker and removes it when the calculation ends, but a copy pulled back by an execution backend's snapshot survives, and `read_progress` kept preferring it: a finished case reported the iteration that snapshot was taken at, for ever. Progress is now read from the log once a case is terminal. A second defect surfaced with it: `extract_last_iteration` did not recognise `TIME STEP NUMBER`, the marker the solver actually announces, and only worked when incidental warnings happened to mention a step
- Selecting a case with no results in the Residuals, Probes or Profiles panel made the panel vanish, taking the case selector with it and leaving no way to pick another one. The controls were rendered inside the same condition as the plot, so an empty result set hid both; they now render whenever the campaign has cases, and the plot area says which of the two is missing. The Probes and Profiles tabs had a second cause: the loader overwrote the flag that enables a tab, which answers "does this campaign have probe data at all", with one scoped to the current selection, so picking a case that had not run disabled the tab itself
- The RESU size column stopped moving. The cached figure was keyed on the results directory's mtime, and some filesystems, WSL2 among them, never bump a directory's mtime when files are created or changed inside it, so the cache never invalidated and the size stayed frozen for the life of the process. It now also expires after ten seconds
- The live residual curve plotted the wrong abscissa. While a run is in progress the points come from the solver log, and each one was numbered by counting the convergence blocks printed so far rather than reading the time step the solver announces. Convergence is printed at the listing frequency, so the Nth block is almost never iteration N: a run at step 5000 showed a curve ending at 44. Affects every run, local ones included; it is simply more visible on a long one
- A case launched on Qarnot could land on a machine with fewer cores than the run asked for. The rank and thread counts only ever reached the solver, inside `DOCKER_CMD`, while the task itself carried no hardware constraint, so Qarnot allocated any available machine and MPI oversubscribed on paid compute. Submission now asks for at least `n` x `nt` cores
- A code_saturne case launched on Qarnot failed in preprocessing with `Mesh file ... not found (no mesh directory given)`. A remote task has one writable directory, so the campaign's shared directories land inside the case rather than beside it, while code_saturne resolves a bare mesh name against `<study>/MESH`. Adapters gain `prepare_remote_case`, called before any backend launch, and code_saturne's adds a `<meshdir>` entry to `setup.xml`; the entry is harmless locally, where the study directory stays a fallback. A task now returns only its results directory, so the shared mesh it was handed is not re-downloaded on every run
- Asking Qarnot for a minimum core count broke every launch with `Some constraints don't exist. Invalid hardware constraints.` Hardware constraints are validated against a per-account catalogue and cannot be invented, so the core constraint is now sent only when the account actually offers it, and dropped otherwise. The node type chosen in the dialog comes from that same catalogue, so it is always sent
- `registry_lock` deadlocked against itself when taken twice in the same thread, freezing the whole process silently and permanently. `REGISTRY_THREAD_LOCK` is an `RLock`, so a nested call passed straight through it, but `flock` applies per file descriptor and the nested call opened a second one, which waited on the first. Launching a case on the Qarnot backend hit exactly this (`_shared_bucket` reads the registry to decide whether the shared mesh still needs uploading) and hung the server before a single request reached the provider. The lock is now genuinely re-entrant within a thread
- Launching a case on an execution backend froze the whole dashboard. The submit uploads the case and talks to a remote API, and it ran inside `registry_transaction`, so every reader of `registry.json` blocked for its full duration: `/api/status` calls `load_registry`, which takes the same lock, and the UI showed nothing changing while the case sat in `PENDING`. The submit now runs with no lock held and writes its outcome back under its own short transaction, the three-phase discipline the synchronisation pass already followed
- Stopping a case on an execution backend that had just finished raised instead of doing nothing. Qarnot refuses to abort a finished task, so `kill_case` returned HTTP 500 on a case that completed between the last status refresh and the button press; a terminal task is now left alone, and the same refusal is tolerated when the task ends mid-call
- `csauto doctor --backend qarnot` reported only that a token was present, so an exhausted bucket or storage quota was discovered as a `QuotaExceeded` in the middle of an upload, after a campaign had been chosen and launched. It now reports the account's bucket count and storage use, and warns before either runs out
- A case whose launch failed outright (no docker or `sbatch` on PATH, an unreadable binary, a rejected `sbatch` submission) stayed `PENDING` forever. `_launch_local` and `_launch_slurm` record `status=FAILED` and then raise to report the failure, but `registry_transaction` skipped its save whenever the body raised, so that write was discarded. The save now runs in a `finally`, which also stops `csauto prepare` from discarding the registry entries of cases it already created on disk when it aborts part way
- A case launched from the web UI stayed `RUNNING` forever when the run crashed before the solver started (a missing mesh, say). Two causes, both fixed: `is_process_alive` reported a zombie as alive, because `os.kill(pid, 0)` succeeds on a process that exited but was never reaped, which is what every run launched by the long-lived server becomes; and `detect_run_outcome` only read `run_solver.log` and `listing`, so a failure that produces neither went undetected even though code_saturne writes an explicit `run_status.failed` marker next to them. Zombies are now reported as dead (and reaped when we are the parent), and the status markers are read when the logs give no verdict
- Opening the solver GUI on a case whose shared dirs are symlinks (the default since `mesh_mode = "symlink"` became the default in 0.4.1) left those symlinks dangling inside the container: `build_gui_command` (docker) and the singularity branch of `build_runtime_gui_command` mounted only the runs dir, unlike their `run` counterparts which also bind the symlink targets. Both now bind them the same way, `MESH` read-only and `POST` writable
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -183,6 +183,7 @@ To re-enable: `csauto enable-telemetry`
- [Web UI guide](docs/web-ui.md)
- [CLI reference](docs/cli.md)
- [Configuration](docs/config.md)
- [Running on Qarnot](docs/qarnot.md)
- [DOE format](docs/doe-format.md)
- [Run lifecycle](docs/run-lifecycle.md)
- [HTTP API](docs/api.md)
Expand Down
173 changes: 173 additions & 0 deletions csauto/backend_sync.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,173 @@
"""One synchronisation pass over the cases a remote backend owns.

Kept out of runner.py, which is already the largest module in the project. The
web server calls this on a timer; `csauto status` calls it once, throttled.

Follows the three-phase pattern `_count_running_cases` uses: snapshot under a
brief lock, do the network work with no lock held, write back under a new lock.
Holding the registry lock across a network round trip would block the CLI and
every other writer for its duration.
"""

from __future__ import annotations

import contextlib
from collections.abc import Callable
from datetime import datetime
from pathlib import Path
from typing import Any

from .backends import ExecutionBackend, get_backend
from .registry import (
STATUS_DONE,
STATUS_FAILED,
STATUS_RUNNING,
load_registry,
mutate_registry,
timestamp_now,
)
from .warn import warn

TERMINAL_STATUSES = (STATUS_DONE, STATUS_FAILED)

# A user may click Refresh as fast as they like, and `csauto status` may be run
# in a loop. Neither should become a provider round trip each time.
SYNC_THROTTLE_S = 10.0


def sync_backend_cases(
runs_dir: Path,
*,
sync_files: bool = True,
backend_factory: Callable[[str], ExecutionBackend] = get_backend,
) -> int:
"""Poll every unfinished backend case once. Returns how many were touched."""
# Phase 1: snapshot the cases a backend owns, under a brief lock.
snapshot: dict[str, dict[str, Any]] = {}
for case_id, record in load_registry(runs_dir).items():
if not str(record.get("backend") or "") or not str(record.get("task_id") or ""):
continue
if record.get("status") in TERMINAL_STATUSES and record.get("results_fetched"):
continue
snapshot[case_id] = dict(record)

# Phase 2: network work, with no lock held.
updates: dict[str, dict[str, Any]] = {}
for case_id, record in snapshot.items():
backend = backend_factory(str(record["backend"]))
updates[case_id] = _observe(runs_dir, case_id, record, backend, str(record["task_id"]), sync_files)

# Phase 3: write back under a new lock, skipping any case relaunched meanwhile.
if updates:

def apply(registry: dict[str, dict[str, Any]]) -> bool:
changed = False
for case_id, fields in updates.items():
current = registry.get(case_id)
if current is None:
continue
if str(current.get("task_id") or "") != str(snapshot[case_id].get("task_id") or ""):
continue
current.update(fields)
current["last_update"] = timestamp_now()
changed = True
return changed

mutate_registry(runs_dir, apply)
return len(snapshot)


def _observe(
runs_dir: Path,
case_id: str,
record: dict[str, Any],
backend: ExecutionBackend,
task_id: str,
sync_files: bool,
) -> dict[str, Any]:
"""Talk to the backend and return the registry fields to write. No lock held."""
case_dir = Path(record.get("path") or (runs_dir / case_id))

try:
state = backend.poll(task_id)
except Exception as exc:
# A transient failure must never change the status: a flaky network
# would otherwise mark a whole campaign as failed.
warn(f"{case_id}: backend poll failed ({exc})")
return {"backend_poll_failures": int(record.get("backend_poll_failures") or 0) + 1}

fields: dict[str, Any] = {"backend_poll_failures": 0}
if state.progress is not None:
fields["backend_progress"] = state.progress
if state.execution_time_s is not None:
fields["backend_execution_time_s"] = state.execution_time_s
if state.running_core_count is not None:
fields["backend_core_count"] = state.running_core_count
_append(case_dir / "csauto.stdout", state.stdout_delta)
_append(case_dir / "csauto.stderr", state.stderr_delta)

if sync_files and state.status == STATUS_RUNNING:
try:
backend.sync(task_id, case_dir)
except Exception as exc:
warn(f"{case_id}: backend sync failed ({exc})")

if state.status not in TERMINAL_STATUSES:
fields["status"] = state.status
return fields

try:
backend.fetch_final(task_id, case_dir)
except Exception as exc:
# The remote task finished but the results are not local yet. DONE must
# mean the results are on disk, so leave the status alone and retry.
warn(f"{case_id}: downloading results failed, will retry ({exc})")
return fields

fields["results_fetched"] = True
fields["status"] = state.status
fields["end_time"] = record.get("end_time") or timestamp_now()
return fields


def _append(path: Path, text: str) -> None:
if not text:
return
try:
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("a", encoding="utf-8") as handle:
handle.write(text)
except OSError:
return


def sync_backend_cases_throttled(
runs_dir: Path,
*,
sync_files: bool = True,
backend_factory: Callable[[str], ExecutionBackend] = get_backend,
) -> bool:
"""Run one pass, at most once every 10 s. Returns whether it ran.

Never raises: the caller is a status command or a Refresh click, and a
provider outage must not fail either. The marker lives under a reserved
`_backend` key at the top of registry.json, which no case id can collide
with: `_resolve_case_id` rejects ids not matching ^[A-Za-z0-9][A-Za-z0-9_.-]*$.
"""
marker = load_registry(runs_dir).get("_backend", {}).get("last_sync")
if marker:
with contextlib.suppress(ValueError):
if (datetime.now() - datetime.fromisoformat(str(marker))).total_seconds() < SYNC_THROTTLE_S:
return False

def stamp(registry: dict[str, dict[str, Any]]) -> bool:
registry.setdefault("_backend", {})["last_sync"] = timestamp_now()
return True

mutate_registry(runs_dir, stamp)
with contextlib.suppress(Exception):
sync_backend_cases(runs_dir, sync_files=sync_files, backend_factory=backend_factory)
return True


__all__ = ["sync_backend_cases", "sync_backend_cases_throttled"]
40 changes: 40 additions & 0 deletions csauto/backends/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
"""Execution backend registry.

The `backend` field on a case selects which service runs it. Implementation
modules are imported lazily so that loading configuration never pulls in an
optional SDK.
"""

from __future__ import annotations

import functools

from .base import BackendState, ExecutionBackend

_BACKEND_NAMES = ("fake", "qarnot")


def available_backends() -> tuple[str, ...]:
return _BACKEND_NAMES


def get_backend(name: str) -> ExecutionBackend:
"""Return the memoized backend instance for a name."""
return _backend_for(str(name or "").strip().lower())


@functools.cache
def _backend_for(normalized: str) -> ExecutionBackend:
if normalized == "fake":
from .fake import FakeBackend

return FakeBackend()
if normalized == "qarnot":
from .qarnot import QarnotBackend

return QarnotBackend()
choices = ", ".join(available_backends())
raise ValueError(f"Unknown backend: {normalized!r}. Choices: {choices}")


__all__ = ["BackendState", "ExecutionBackend", "available_backends", "get_backend"]
Loading
Loading