Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions integrations/agentbricks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1029,3 +1029,20 @@ sync/streaming/background transport selector is manual.
Developing Agent Bricks CLI (`agentbricks`), AgentKit, the runtime, and templates - plus the local dev loop and how to
test unreleased changes on `agentbricks dev` and `agentbricks deploy`, is covered in
[CONTRIBUTING.md](CONTRIBUTING.md).

## Project inventory, cleanup and evaluations

`agentbricks --profile <profile> status` shows project bindings without modifying resources.
Add `--verify` for read-only workspace checks; saved configuration is clearly distinguished from
verified resource availability. `agentbricks --profile <profile> cleanup` previews retained and
removable resources. `cleanup --apply` asks before deleting project-created Apps and their
owner-validated managed Runtime Stores. Shared stores, experiments, tools and source files are
retained; see the [command reference](cli.md#agentbricks-cleanup) for the ownership boundary.

Both managed-runtime scaffolds include a small extendable evaluation dataset. With the agent
running, `uv run python evals/run.py` invokes that actual agent and records case results and
aggregate checks in MLflow. The scaffold's `evals/README.md` explains how to extend the dataset,
inspect failures and evaluate a deployment. These are smoke checks, not domain-quality certification.

The [project overview design](docs/project-overview-design.md) scopes UI resource, cost, evaluation
and deployed-version summaries, including required data sources and unsupported states.
73 changes: 73 additions & 0 deletions integrations/agentbricks/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,8 @@ These options apply to every command. Pass them before the command name, for exa
| [`logout`](#agentbricks-logout) | Forget the saved default profile |
| [`init`](#agentbricks-init) | Scaffold a new agent project |
| [`doctor`](#agentbricks-doctor) | Check an existing agent's Agent Bricks onboarding |
| [`status`](#agentbricks-status) | Read project bindings and optionally verify resource availability |
| [`cleanup`](#agentbricks-cleanup) | Preview or delete project-created deployments while retaining shared data |
| [`dev`](#agentbricks-dev) | Run the agent locally with a chat UI |
| [`memory`](#agentbricks-memory) | Manage an agent's long-term memory |
| [`mcp`](#agentbricks-mcp) | Discover managed MCP services |
Expand Down Expand Up @@ -1426,3 +1428,74 @@ _Options_
| Option | Values | Default | Required | Description |
| --- | --- | --- | --- | --- |
| `--source <SOURCE>` | path | `.` | no | Agent project containing agent.toml. |

## `agentbricks status`

Show project bindings and resources; configuration alone is not verified readiness.

```sh
agentbricks --profile <profile> status [--source DIRECTORY] [--verify]
agentbricks --profile <profile> --output json status --verify
```

| Option | Default | Description |
| --- | --- | --- |
| `--source DIRECTORY` | `.` | Project directory containing `agent.toml`. |
| `--verify` | off | Check resource availability using read-only workspace APIs. Requires a selected profile. |

Without `--verify`, status reads only local configuration and provisioning receipts. It does not
initialize authentication, refresh tokens, create stores, or write project files. The workspace host
comes from the selected profile; no profile means an unresolved workspace. With `--verify`, status
uses the selected profile's actual workspace and reports resource identifiers, URLs when available,
App state, missing resources and individual check errors without losing successful results.
`accessible` means the caller can read the resource; it does not promise the deployed App can use
it, or that model/tool invocation has succeeded. Tool bindings remain unverified. JSON contains a
versioned inventory with per-resource ownership, verification, error and next action.

## `agentbricks cleanup`

Preview cleanup of project-created deployments; retain shared stores, traces and tools.

```sh
agentbricks --profile <profile> cleanup [--source DIRECTORY]
agentbricks --profile <profile> cleanup --apply [--yes] [--source DIRECTORY]
```

| Option | Default | Description |
| --- | --- | --- |
| `--source DIRECTORY` | `.` | Project directory containing `agent.toml`. |
| `--apply` | off | Apply the previewed cleanup; default is preview only. |
| `--yes` (`-y`) | off | Skip the confirmation prompt. Requires `--apply`. |

A selected profile scopes both preview and apply to one workspace. The default preview uses only
local files. Apply resolves the actual workspace, displays its exact targets, then asks for
confirmation. `--output json` provides the final per-resource plan/results on stdout; the apply
preview is written to stderr. Failed resources cause exit status 1 and can be retried.

Deploy writes workspace-scoped creation/reuse receipts to `.agentbricks/resources.json`, including
resources created before a later deploy step fails. Keep this local file to enable safe cleanup.
Only a project-created App whose current service-principal identity still matches its receipt can
be removed. Its managed Runtime Store is removed first, after verifying its app owner; failure
retains the App. Successful deletes and partial failures are persisted after each resource.

Existing/adopted Apps and resources with no receipt are retained. Shared-capable memory/session
stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files are
always retained, even when this project created them. Current APIs cannot prove exclusive use.
If the App is already missing, any residual Runtime Store is retained for manual owner inspection.
Cleanup does not unbind resources or erase source/evaluation history.

## Starter evaluations

The OpenAI and LangGraph Agent Bricks server templates include `evals/cases.jsonl`, `evals/run.py`
and extension instructions. Start the actual project agent with `agentbricks dev`, then run:

```sh
uv run python evals/run.py
uv run python evals/run.py --app <deployment> --profile <oauth-profile>
```

The runner calls `/api/invocations`, uses fresh sessions, preserves the project model/tools, and
records case inputs, final answers, errors, scorer feedback and aggregate results in MLflow.
The default tracking URI is `sqlite:///.agentbricks/evaluations.db`. See the scaffold's
`evals/README.md` for all options, results UI and workspace tracking. A failed invocation or
expectation exits nonzero. Custom HTTP server templates require their own evaluator contract.
54 changes: 54 additions & 0 deletions integrations/agentbricks/docs/project-overview-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# Project overview: initial scope and follow-ups

The CLI now exposes a truthful resource inventory with `agentbricks status`. A UI overview should
reuse this read-only model rather than infer readiness from a saved binding. This document scopes
the requested resource, cost, evaluation and deployed-version panels. It does not imply that the
panels have shipped.

## First implementation slice: resources and links

Show the project framework, selected workspace, declared bindings and resolved resource identifiers
with links where the platform exposes a stable destination. Distinguish **configured**, **checking**,
**accessible**, **missing**, and **check failed**. An accessible resource does not prove the running
app's identity can invoke it; runtime validation belongs to a separate check. Show the last check
time, request identity, workspace and action for each unavailable resource. No automatic provisioning
or deletion occurs on page load. Refresh is a read operation.

The initial slice is a resources card and links to MLflow and Databricks Apps, with the CLI status
schema as its contract. The deployed UI needs a server route that exposes an allowlist of this
information to authorized project maintainers. It must not expose local filesystem paths,
credentials, arbitrary app environment variables, or administrative actions to chat end users.

| Panel | Truthful source | Missing/unavailable states | Scope and dependencies |
| --- | --- | --- | --- |
| Resources | `agent.toml` intent; workspace-specific provisioning receipts; read-only Apps, Session Store, Memory Store, and MLflow experiment APIs | Unverified binding; resource absent; API unavailable; insufficient permission; receipt unavailable for older projects | First slice. Resource IDs and supplied URLs; tool bindings remain unverified until invoked. Receipts are a cleanup aid, not an access-control authority. |
| Cost | Approved billing/usage system tables with explicit workspace/resource attribution and published price assumptions; trace token usage for diagnostic counts only | No billing permission; attribution unavailable; incomplete time window; delayed usage; unknown price | Follow-up owned with billing. Do not label token counts or partial traces as dollar cost. Show window, currency, source, coverage and refresh time before any total. |
| Evaluations | MLflow evaluation run ID, dataset identity, scorer names/versions, per-case outcomes and aggregate metrics from the starter or project-specific suite | No runs yet; run failed; different dataset/scorers; deleted or inaccessible experiment | First slice links to the existing MLflow result page. Inline comparisons require the same dataset/scorer version and an authorized read API; never imply the smoke dataset certifies production quality. |
| Version | Apps deployment metadata (deployment ID/state/source path), plus explicitly captured source revision and installed package/template version | No deployment; deployment failed; local changes; source revision not recorded; old receipt without version | First slice links to Apps deployment details. Follow-up captures commit SHA, dirty-tree marker and artifact identity at deploy; it must not substitute the local Git HEAD for the deployed revision. |

## Cleanup boundary

`agentbricks cleanup` previews resources in the explicitly selected workspace. `--apply` requires
confirmation (or `--yes` for automation). Only Apps with a local creation receipt are candidates;
the current remote service-principal identity must match before deletion. Managed Runtime Store
deletion also validates the API's app owner. A failed store delete retains the App for retry.
Results are saved after each resource, so partial failures can be retried independently.

Stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files
are retained, including stores this project originally created. The current APIs cannot prove
these resources have no other consumers. The preview identifies those retained resources rather
than silently destroying shared data. Missing receipts, copied/adopted Apps and changed App identity
also fail closed. A future explicit store-deletion workflow requires service-supported ownership
and consumer checks; a display-name prefix is insufficient.

## Review checklist for the UI follow-up

- Validate fresh, partly configured, deployed and permission-denied states without provisioning.
- Establish maintainer authorization before exposing administrative project information.
- Verify each link and data field against the selected workspace and request identity.
- Show unavailable cost/version information explicitly; do not fill gaps with estimates.
- Keep chat execution and its errors usable if the overview API fails.

Source feedback: Chang Shi Lim, *Agentbricks_CLI_Feedback*, Wish List items 1–3 (UGW resources, Cost,
Evals, Version); Fabian Nobis, *Production Grade Document Chatbot*, paragraph beginning “One thing
I’m missing is scaffolding for an eval set.” Tracked in ML-70357.
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
from databricks_agentbricks.cli.help import configure_help
from databricks_agentbricks.cli.init import init
from databricks_agentbricks.cli.memory import memory
from databricks_agentbricks.cli.project import cleanup, status
from databricks_agentbricks.cli.sessions import sessions
from databricks_agentbricks.cli.tools import tools
from databricks_agentbricks.cli.tracing import tracing
Expand Down Expand Up @@ -90,6 +91,8 @@ def agentbricks(ctx: click.Context, profile: Optional[str], output: str) -> None
agentbricks.add_command(logout)
agentbricks.add_command(init)
agentbricks.add_command(doctor)
agentbricks.add_command(status)
agentbricks.add_command(cleanup)
agentbricks.add_command(dev)
agentbricks.add_command(memory)
agentbricks.add_command(sessions)
Expand Down
75 changes: 70 additions & 5 deletions integrations/agentbricks/src/databricks_agentbricks/cli/deploy.py
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,7 @@
from databricks_agentbricks.project_config import (
require_managed_tool_support,
)
from databricks_agentbricks.project_resources import record_resource
from databricks_agentbricks.project_types import AgentServer
from databricks_agentbricks.render import field
from databricks_agentbricks.tool_access import (
Expand Down Expand Up @@ -116,6 +117,25 @@ def _app_url(name: str, profile: Optional[str]) -> Optional[str]:
return None


def _record_deployment_receipt(source, name, obj, *, created, managed_runtime) -> None:
principal = _app_service_principal(name, obj.profile)
if principal:
record_resource(
source,
host=obj.client().host,
kind="deployment",
name=name,
resource_id=principal,
created=created,
managed_runtime=managed_runtime,
)
else:
click.echo(
"Could not record the app identity; project cleanup will retain this deployment.",
err=True,
)


def _app_compute_state(name: str, profile: Optional[str]) -> Optional[str]:
"""The app's compute state (e.g. RUNNING), or None if it can't be read."""
result = _databricks(["apps", "get", name, "-o", "json"], profile, capture=True, check=False)
Expand Down Expand Up @@ -166,11 +186,11 @@ def _prefixed_name(name: str) -> str:
return name if name.startswith(_DEPLOYMENT_PREFIX) else f"{_DEPLOYMENT_PREFIX}{name}"


def _confirm_destroy(target: str, *, assume_yes: bool) -> None:
def _confirm_destroy(target: str, *, assume_yes: bool, err: bool = False) -> None:
"""Prompt before a destructive deployment op; --yes/-y skips it (for scripts)."""
if assume_yes:
return
if not click.confirm(f"{target}? This cannot be undone.", default=False):
if not click.confirm(f"{target}? This cannot be undone.", default=False, err=err):
raise click.Abort()


Expand Down Expand Up @@ -367,7 +387,11 @@ def _resolve_deployment_name(project, name: Optional[str]) -> str:


def _reconcile_declared_stores(
memory_store: Optional[str], session_store: Optional[str], client
memory_store: Optional[str],
session_store: Optional[str],
client,
*,
source: pathlib.Path | None = None,
) -> Optional[str]:
"""Create any store DECLARED in agent.toml that doesn't exist yet; return the memory store's id.

Expand All @@ -382,11 +406,29 @@ def _reconcile_declared_stores(
with render.status(f"Reconciling memory store '{memory_store}'…"):
resolved, created = _ensure_memory_store(client, memory_store)
memory_store_id = (field(resolved, "name") or "").split("/", 1)[-1] or None
if source is not None and memory_store_id:
record_resource(
source,
host=client.host,
kind="memory_store",
name=memory_store,
resource_id=memory_store_id,
created=created,
)
if created:
render.console().print(f"[green]✓[/] Created memory store {memory_store!r}")
if session_store:
with render.status(f"Reconciling session store '{session_store}'…"):
_, created = _ensure_session_store(client, session_store)
resolved, created = _ensure_session_store(client, session_store)
if source is not None:
record_resource(
source,
host=client.host,
kind="session_store",
name=session_store,
resource_id=field(resolved, "name") or session_store,
created=created,
)
if created:
render.console().print(f"[green]✓[/] Created session store {session_store!r}")
return memory_store_id
Expand Down Expand Up @@ -586,6 +628,15 @@ def deploy(
if deployment_exists is None:
deployment_exists = user_scope_plan.existing_scopes is not None
apply_app_user_scope_update(user_scope_plan, instances=instances)
_record_deployment_receipt(
source_dir,
name,
obj,
created=not deployment_exists,
managed_runtime=bool(
project and project.server == AgentServer.AGENTBRICKS and _USE_MANAGED_RUNTIME_STORE
),
)
# Persist the base name so a later `agentbricks deploy` (no NAME) resolves to the same app.
if project is not None and project.set_deployment_name(base_name):
project.write()
Expand All @@ -596,7 +647,9 @@ def deploy(
# 1. Reconcile the stores DECLARED in agent.toml: create any that don't exist yet. `agentbricks deploy`
# is the only reconcile-to-cloud verb; agent.toml is the source of truth and is never rewritten.
memory_store, session_store, _ = resource_bindings(source_dir)
memory_store_id = _reconcile_declared_stores(memory_store, session_store, client)
memory_store_id = _reconcile_declared_stores(
memory_store, session_store, client, source=source_dir
)

# 2. Provision tracing when bound (`agentbricks init` binds a default experiment): get-or-create the
# experiment NAME from agent.toml and wire the two env vars the runtime reads. Resolved by name,
Expand Down Expand Up @@ -696,6 +749,18 @@ def deploy(
)
old, new = _AGENT_COMPUTE_OUTPUT
click.echo((result.stdout or "").replace(old, new), nl=False)
# Record creation before later provisioning/upload steps can fail. An existing app is adopted,
# even when its name matches this project; cleanup must never infer ownership from a name.
if user_scope_plan is None:
_record_deployment_receipt(
source_dir,
name,
obj,
created=not deployment_exists,
managed_runtime=bool(
project and project.server == AgentServer.AGENTBRICKS and use_managed_runtime_store
),
)
# `apps deploy` requires the app's compute to be ACTIVE — a just-created app may still be
# starting, and an existing one may be STOPPED — so wait either way. Returns immediately when
with render.progress("Waiting for agent compute to start (this can take a few minutes)…"):
Expand Down
18 changes: 16 additions & 2 deletions integrations/agentbricks/src/databricks_agentbricks/cli/help.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,9 +16,9 @@
# getting-started path. Any command missing here still lists under "Other commands" (see
# `_group.AgentBricksGroup`).
_COMMAND_SECTIONS: tuple[tuple[str, tuple[str, ...]], ...] = (
("SETUP", ("login", "logout", "init", "doctor")),
("SETUP", ("login", "logout", "init", "doctor", "status")),
("DEVELOP", ("dev", "tools", "memory", "sessions", "tracing")),
("SHIP", ("deploy", "deployments")),
("SHIP", ("deploy", "deployments", "cleanup")),
)

# Each example is either a bare command, or a (command, comment) pair. The comment is a short gloss
Expand Down Expand Up @@ -46,6 +46,20 @@
("doctor",): (
("agentbricks doctor .", "check an existing repository's Agent Bricks onboarding"),
),
("status",): (
("agentbricks --profile <profile> status", "read project bindings without workspace calls"),
(
"agentbricks --profile <profile> status --verify",
"check resource availability without provisioning",
),
),
("cleanup",): (
("agentbricks --profile <profile> cleanup", "preview deletions and retained resources"),
(
"agentbricks --profile <profile> cleanup --apply",
"confirm deletion of project-created deployments",
),
),
("dev",): (("agentbricks dev", "run the agent locally with a chat UI"),),
("memory",): (
("agentbricks memory stores create --display-name agent-memory", "create a memory store"),
Expand Down
Loading
Loading