diff --git a/integrations/agentbricks/README.md b/integrations/agentbricks/README.md index 624bc3e1c..55b8c29fe 100644 --- a/integrations/agentbricks/README.md +++ b/integrations/agentbricks/README.md @@ -1029,3 +1029,20 @@ sync/streaming/background transport selector is manual. Developing Agent Bricks CLI (`agentbricks`), AgentKit, the runtime, and templates - plus the local dev loop and how to test unreleased changes on `agentbricks dev` and `agentbricks deploy`, is covered in [CONTRIBUTING.md](CONTRIBUTING.md). + +## Project inventory, cleanup and evaluations + +`agentbricks --profile status` shows project bindings without modifying resources. +Add `--verify` for read-only workspace checks; saved configuration is clearly distinguished from +verified resource availability. `agentbricks --profile cleanup` previews retained and +removable resources. `cleanup --apply` asks before deleting project-created Apps and their +owner-validated managed Runtime Stores. Shared stores, experiments, tools and source files are +retained; see the [command reference](cli.md#agentbricks-cleanup) for the ownership boundary. + +Both managed-runtime scaffolds include a small extendable evaluation dataset. With the agent +running, `uv run python evals/run.py` invokes that actual agent and records case results and +aggregate checks in MLflow. The scaffold's `evals/README.md` explains how to extend the dataset, +inspect failures and evaluate a deployment. These are smoke checks, not domain-quality certification. + +The [project overview design](docs/project-overview-design.md) scopes UI resource, cost, evaluation +and deployed-version summaries, including required data sources and unsupported states. diff --git a/integrations/agentbricks/cli.md b/integrations/agentbricks/cli.md index 1cde69db2..70765e034 100644 --- a/integrations/agentbricks/cli.md +++ b/integrations/agentbricks/cli.md @@ -61,6 +61,8 @@ These options apply to every command. Pass them before the command name, for exa | [`logout`](#agentbricks-logout) | Forget the saved default profile | | [`init`](#agentbricks-init) | Scaffold a new agent project | | [`doctor`](#agentbricks-doctor) | Check an existing agent's Agent Bricks onboarding | +| [`status`](#agentbricks-status) | Read project bindings and optionally verify resource availability | +| [`cleanup`](#agentbricks-cleanup) | Preview or delete project-created deployments while retaining shared data | | [`dev`](#agentbricks-dev) | Run the agent locally with a chat UI | | [`memory`](#agentbricks-memory) | Manage an agent's long-term memory | | [`mcp`](#agentbricks-mcp) | Discover managed MCP services | @@ -1426,3 +1428,74 @@ _Options_ | Option | Values | Default | Required | Description | | --- | --- | --- | --- | --- | | `--source ` | path | `.` | no | Agent project containing agent.toml. | + +## `agentbricks status` + +Show project bindings and resources; configuration alone is not verified readiness. + +```sh +agentbricks --profile status [--source DIRECTORY] [--verify] +agentbricks --profile --output json status --verify +``` + +| Option | Default | Description | +| --- | --- | --- | +| `--source DIRECTORY` | `.` | Project directory containing `agent.toml`. | +| `--verify` | off | Check resource availability using read-only workspace APIs. Requires a selected profile. | + +Without `--verify`, status reads only local configuration and provisioning receipts. It does not +initialize authentication, refresh tokens, create stores, or write project files. The workspace host +comes from the selected profile; no profile means an unresolved workspace. With `--verify`, status +uses the selected profile's actual workspace and reports resource identifiers, URLs when available, +App state, missing resources and individual check errors without losing successful results. +`accessible` means the caller can read the resource; it does not promise the deployed App can use +it, or that model/tool invocation has succeeded. Tool bindings remain unverified. JSON contains a +versioned inventory with per-resource ownership, verification, error and next action. + +## `agentbricks cleanup` + +Preview cleanup of project-created deployments; retain shared stores, traces and tools. + +```sh +agentbricks --profile cleanup [--source DIRECTORY] +agentbricks --profile cleanup --apply [--yes] [--source DIRECTORY] +``` + +| Option | Default | Description | +| --- | --- | --- | +| `--source DIRECTORY` | `.` | Project directory containing `agent.toml`. | +| `--apply` | off | Apply the previewed cleanup; default is preview only. | +| `--yes` (`-y`) | off | Skip the confirmation prompt. Requires `--apply`. | + +A selected profile scopes both preview and apply to one workspace. The default preview uses only +local files. Apply resolves the actual workspace, displays its exact targets, then asks for +confirmation. `--output json` provides the final per-resource plan/results on stdout; the apply +preview is written to stderr. Failed resources cause exit status 1 and can be retried. + +Deploy writes workspace-scoped creation/reuse receipts to `.agentbricks/resources.json`, including +resources created before a later deploy step fails. Keep this local file to enable safe cleanup. +Only a project-created App whose current service-principal identity still matches its receipt can +be removed. Its managed Runtime Store is removed first, after verifying its app owner; failure +retains the App. Successful deletes and partial failures are persisted after each resource. + +Existing/adopted Apps and resources with no receipt are retained. Shared-capable memory/session +stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files are +always retained, even when this project created them. Current APIs cannot prove exclusive use. +If the App is already missing, any residual Runtime Store is retained for manual owner inspection. +Cleanup does not unbind resources or erase source/evaluation history. + +## Starter evaluations + +The OpenAI and LangGraph Agent Bricks server templates include `evals/cases.jsonl`, `evals/run.py` +and extension instructions. Start the actual project agent with `agentbricks dev`, then run: + +```sh +uv run python evals/run.py +uv run python evals/run.py --app --profile +``` + +The runner calls `/api/invocations`, uses fresh sessions, preserves the project model/tools, and +records case inputs, final answers, errors, scorer feedback and aggregate results in MLflow. +The default tracking URI is `sqlite:///.agentbricks/evaluations.db`. See the scaffold's +`evals/README.md` for all options, results UI and workspace tracking. A failed invocation or +expectation exits nonzero. Custom HTTP server templates require their own evaluator contract. diff --git a/integrations/agentbricks/docs/project-overview-design.md b/integrations/agentbricks/docs/project-overview-design.md new file mode 100644 index 000000000..2578897a1 --- /dev/null +++ b/integrations/agentbricks/docs/project-overview-design.md @@ -0,0 +1,54 @@ +# Project overview: initial scope and follow-ups + +The CLI now exposes a truthful resource inventory with `agentbricks status`. A UI overview should +reuse this read-only model rather than infer readiness from a saved binding. This document scopes +the requested resource, cost, evaluation and deployed-version panels. It does not imply that the +panels have shipped. + +## First implementation slice: resources and links + +Show the project framework, selected workspace, declared bindings and resolved resource identifiers +with links where the platform exposes a stable destination. Distinguish **configured**, **checking**, +**accessible**, **missing**, and **check failed**. An accessible resource does not prove the running +app's identity can invoke it; runtime validation belongs to a separate check. Show the last check +time, request identity, workspace and action for each unavailable resource. No automatic provisioning +or deletion occurs on page load. Refresh is a read operation. + +The initial slice is a resources card and links to MLflow and Databricks Apps, with the CLI status +schema as its contract. The deployed UI needs a server route that exposes an allowlist of this +information to authorized project maintainers. It must not expose local filesystem paths, +credentials, arbitrary app environment variables, or administrative actions to chat end users. + +| Panel | Truthful source | Missing/unavailable states | Scope and dependencies | +| --- | --- | --- | --- | +| Resources | `agent.toml` intent; workspace-specific provisioning receipts; read-only Apps, Session Store, Memory Store, and MLflow experiment APIs | Unverified binding; resource absent; API unavailable; insufficient permission; receipt unavailable for older projects | First slice. Resource IDs and supplied URLs; tool bindings remain unverified until invoked. Receipts are a cleanup aid, not an access-control authority. | +| Cost | Approved billing/usage system tables with explicit workspace/resource attribution and published price assumptions; trace token usage for diagnostic counts only | No billing permission; attribution unavailable; incomplete time window; delayed usage; unknown price | Follow-up owned with billing. Do not label token counts or partial traces as dollar cost. Show window, currency, source, coverage and refresh time before any total. | +| Evaluations | MLflow evaluation run ID, dataset identity, scorer names/versions, per-case outcomes and aggregate metrics from the starter or project-specific suite | No runs yet; run failed; different dataset/scorers; deleted or inaccessible experiment | First slice links to the existing MLflow result page. Inline comparisons require the same dataset/scorer version and an authorized read API; never imply the smoke dataset certifies production quality. | +| Version | Apps deployment metadata (deployment ID/state/source path), plus explicitly captured source revision and installed package/template version | No deployment; deployment failed; local changes; source revision not recorded; old receipt without version | First slice links to Apps deployment details. Follow-up captures commit SHA, dirty-tree marker and artifact identity at deploy; it must not substitute the local Git HEAD for the deployed revision. | + +## Cleanup boundary + +`agentbricks cleanup` previews resources in the explicitly selected workspace. `--apply` requires +confirmation (or `--yes` for automation). Only Apps with a local creation receipt are candidates; +the current remote service-principal identity must match before deletion. Managed Runtime Store +deletion also validates the API's app owner. A failed store delete retains the App for retry. +Results are saved after each resource, so partial failures can be retried independently. + +Stores, experiments, tools, workspace source folders, legacy Lakebase projects and local files +are retained, including stores this project originally created. The current APIs cannot prove +these resources have no other consumers. The preview identifies those retained resources rather +than silently destroying shared data. Missing receipts, copied/adopted Apps and changed App identity +also fail closed. A future explicit store-deletion workflow requires service-supported ownership +and consumer checks; a display-name prefix is insufficient. + +## Review checklist for the UI follow-up + +- Validate fresh, partly configured, deployed and permission-denied states without provisioning. +- Establish maintainer authorization before exposing administrative project information. +- Verify each link and data field against the selected workspace and request identity. +- Show unavailable cost/version information explicitly; do not fill gaps with estimates. +- Keep chat execution and its errors usable if the overview API fails. + +Source feedback: Chang Shi Lim, *Agentbricks_CLI_Feedback*, Wish List items 1–3 (UGW resources, Cost, +Evals, Version); Fabian Nobis, *Production Grade Document Chatbot*, paragraph beginning “One thing +I’m missing is scaffolding for an eval set.” Tracked in ML-70357. diff --git a/integrations/agentbricks/src/databricks_agentbricks/cli/app.py b/integrations/agentbricks/src/databricks_agentbricks/cli/app.py index 0295bde83..e5817821c 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/cli/app.py +++ b/integrations/agentbricks/src/databricks_agentbricks/cli/app.py @@ -19,6 +19,7 @@ from databricks_agentbricks.cli.help import configure_help from databricks_agentbricks.cli.init import init from databricks_agentbricks.cli.memory import memory +from databricks_agentbricks.cli.project import cleanup, status from databricks_agentbricks.cli.sessions import sessions from databricks_agentbricks.cli.tools import tools from databricks_agentbricks.cli.tracing import tracing @@ -90,6 +91,8 @@ def agentbricks(ctx: click.Context, profile: Optional[str], output: str) -> None agentbricks.add_command(logout) agentbricks.add_command(init) agentbricks.add_command(doctor) +agentbricks.add_command(status) +agentbricks.add_command(cleanup) agentbricks.add_command(dev) agentbricks.add_command(memory) agentbricks.add_command(sessions) diff --git a/integrations/agentbricks/src/databricks_agentbricks/cli/deploy.py b/integrations/agentbricks/src/databricks_agentbricks/cli/deploy.py index e77cc5479..395b985b1 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/cli/deploy.py +++ b/integrations/agentbricks/src/databricks_agentbricks/cli/deploy.py @@ -53,6 +53,7 @@ from databricks_agentbricks.project_config import ( require_managed_tool_support, ) +from databricks_agentbricks.project_resources import record_resource from databricks_agentbricks.project_types import AgentServer from databricks_agentbricks.render import field from databricks_agentbricks.tool_access import ( @@ -116,6 +117,25 @@ def _app_url(name: str, profile: Optional[str]) -> Optional[str]: return None +def _record_deployment_receipt(source, name, obj, *, created, managed_runtime) -> None: + principal = _app_service_principal(name, obj.profile) + if principal: + record_resource( + source, + host=obj.client().host, + kind="deployment", + name=name, + resource_id=principal, + created=created, + managed_runtime=managed_runtime, + ) + else: + click.echo( + "Could not record the app identity; project cleanup will retain this deployment.", + err=True, + ) + + def _app_compute_state(name: str, profile: Optional[str]) -> Optional[str]: """The app's compute state (e.g. RUNNING), or None if it can't be read.""" result = _databricks(["apps", "get", name, "-o", "json"], profile, capture=True, check=False) @@ -166,11 +186,11 @@ def _prefixed_name(name: str) -> str: return name if name.startswith(_DEPLOYMENT_PREFIX) else f"{_DEPLOYMENT_PREFIX}{name}" -def _confirm_destroy(target: str, *, assume_yes: bool) -> None: +def _confirm_destroy(target: str, *, assume_yes: bool, err: bool = False) -> None: """Prompt before a destructive deployment op; --yes/-y skips it (for scripts).""" if assume_yes: return - if not click.confirm(f"{target}? This cannot be undone.", default=False): + if not click.confirm(f"{target}? This cannot be undone.", default=False, err=err): raise click.Abort() @@ -367,7 +387,11 @@ def _resolve_deployment_name(project, name: Optional[str]) -> str: def _reconcile_declared_stores( - memory_store: Optional[str], session_store: Optional[str], client + memory_store: Optional[str], + session_store: Optional[str], + client, + *, + source: pathlib.Path | None = None, ) -> Optional[str]: """Create any store DECLARED in agent.toml that doesn't exist yet; return the memory store's id. @@ -382,11 +406,29 @@ def _reconcile_declared_stores( with render.status(f"Reconciling memory store '{memory_store}'…"): resolved, created = _ensure_memory_store(client, memory_store) memory_store_id = (field(resolved, "name") or "").split("/", 1)[-1] or None + if source is not None and memory_store_id: + record_resource( + source, + host=client.host, + kind="memory_store", + name=memory_store, + resource_id=memory_store_id, + created=created, + ) if created: render.console().print(f"[green]✓[/] Created memory store {memory_store!r}") if session_store: with render.status(f"Reconciling session store '{session_store}'…"): - _, created = _ensure_session_store(client, session_store) + resolved, created = _ensure_session_store(client, session_store) + if source is not None: + record_resource( + source, + host=client.host, + kind="session_store", + name=session_store, + resource_id=field(resolved, "name") or session_store, + created=created, + ) if created: render.console().print(f"[green]✓[/] Created session store {session_store!r}") return memory_store_id @@ -586,6 +628,15 @@ def deploy( if deployment_exists is None: deployment_exists = user_scope_plan.existing_scopes is not None apply_app_user_scope_update(user_scope_plan, instances=instances) + _record_deployment_receipt( + source_dir, + name, + obj, + created=not deployment_exists, + managed_runtime=bool( + project and project.server == AgentServer.AGENTBRICKS and _USE_MANAGED_RUNTIME_STORE + ), + ) # Persist the base name so a later `agentbricks deploy` (no NAME) resolves to the same app. if project is not None and project.set_deployment_name(base_name): project.write() @@ -596,7 +647,9 @@ def deploy( # 1. Reconcile the stores DECLARED in agent.toml: create any that don't exist yet. `agentbricks deploy` # is the only reconcile-to-cloud verb; agent.toml is the source of truth and is never rewritten. memory_store, session_store, _ = resource_bindings(source_dir) - memory_store_id = _reconcile_declared_stores(memory_store, session_store, client) + memory_store_id = _reconcile_declared_stores( + memory_store, session_store, client, source=source_dir + ) # 2. Provision tracing when bound (`agentbricks init` binds a default experiment): get-or-create the # experiment NAME from agent.toml and wire the two env vars the runtime reads. Resolved by name, @@ -696,6 +749,18 @@ def deploy( ) old, new = _AGENT_COMPUTE_OUTPUT click.echo((result.stdout or "").replace(old, new), nl=False) + # Record creation before later provisioning/upload steps can fail. An existing app is adopted, + # even when its name matches this project; cleanup must never infer ownership from a name. + if user_scope_plan is None: + _record_deployment_receipt( + source_dir, + name, + obj, + created=not deployment_exists, + managed_runtime=bool( + project and project.server == AgentServer.AGENTBRICKS and use_managed_runtime_store + ), + ) # `apps deploy` requires the app's compute to be ACTIVE — a just-created app may still be # starting, and an existing one may be STOPPED — so wait either way. Returns immediately when with render.progress("Waiting for agent compute to start (this can take a few minutes)…"): diff --git a/integrations/agentbricks/src/databricks_agentbricks/cli/help.py b/integrations/agentbricks/src/databricks_agentbricks/cli/help.py index 5d05746cb..d72096d4b 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/cli/help.py +++ b/integrations/agentbricks/src/databricks_agentbricks/cli/help.py @@ -16,9 +16,9 @@ # getting-started path. Any command missing here still lists under "Other commands" (see # `_group.AgentBricksGroup`). _COMMAND_SECTIONS: tuple[tuple[str, tuple[str, ...]], ...] = ( - ("SETUP", ("login", "logout", "init", "doctor")), + ("SETUP", ("login", "logout", "init", "doctor", "status")), ("DEVELOP", ("dev", "tools", "memory", "sessions", "tracing")), - ("SHIP", ("deploy", "deployments")), + ("SHIP", ("deploy", "deployments", "cleanup")), ) # Each example is either a bare command, or a (command, comment) pair. The comment is a short gloss @@ -46,6 +46,20 @@ ("doctor",): ( ("agentbricks doctor .", "check an existing repository's Agent Bricks onboarding"), ), + ("status",): ( + ("agentbricks --profile status", "read project bindings without workspace calls"), + ( + "agentbricks --profile status --verify", + "check resource availability without provisioning", + ), + ), + ("cleanup",): ( + ("agentbricks --profile cleanup", "preview deletions and retained resources"), + ( + "agentbricks --profile cleanup --apply", + "confirm deletion of project-created deployments", + ), + ), ("dev",): (("agentbricks dev", "run the agent locally with a chat UI"),), ("memory",): ( ("agentbricks memory stores create --display-name agent-memory", "create a memory store"), diff --git a/integrations/agentbricks/src/databricks_agentbricks/cli/init.py b/integrations/agentbricks/src/databricks_agentbricks/cli/init.py index bd6ef71c1..e2b5b178f 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/cli/init.py +++ b/integrations/agentbricks/src/databricks_agentbricks/cli/init.py @@ -107,6 +107,13 @@ def _copy_packaged_template( shutil.copytree( str(src), dest, dirs_exist_ok=index > 0, ignore=shutil.ignore_patterns("__pycache__") ) + if name in {"agent-openai", "agent-langgraph"}: + # A scaffold may install the prior released runtime. Copy the evaluator as standalone + # project code rather than importing a new module absent from that release. + evaluator = resources.files("databricks_agentbricks").joinpath("evaluation.py") + (dest / "evals" / "run.py").write_text( + evaluator.read_text(encoding="utf-8"), encoding="utf-8" + ) if any(overlay.startswith("ui/") for overlay in overlay_names): # UI templates must work with the released runtime installed by the scaffold. Keep # schema discovery with its matching UI rather than requiring a new SDK signature. diff --git a/integrations/agentbricks/src/databricks_agentbricks/cli/project.py b/integrations/agentbricks/src/databricks_agentbricks/cli/project.py new file mode 100644 index 000000000..e7b3e0903 --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/cli/project.py @@ -0,0 +1,344 @@ +"""Read-only project inventory and conservative project cleanup.""" + +from __future__ import annotations + +import configparser +import os +import pathlib +from typing import Any +from urllib.parse import quote + +import click +from databricks.sdk.errors import NotFound + +from databricks_agentbricks import lakebase_runtime_store, render +from databricks_agentbricks.agent_project import AgentProject +from databricks_agentbricks.cli.deploy import ( + _confirm_destroy, + _prefixed_name, + _resolve_memory_store, +) +from databricks_agentbricks.cli.tracing import _mlflow, _set_tracking_uri, experiment_url +from databricks_agentbricks.errors import AgentCliError +from databricks_agentbricks.project_resources import load_receipts, save_receipts + + +def _source_option(function): + return click.option( + "--source", + default=".", + type=click.Path(exists=True, file_okay=False, path_type=pathlib.Path), + help="Project directory containing agent.toml (default: current directory).", + )(function) + + +def _configured_host(profile: str | None) -> str | None: + """Read only the chosen profile's host, without initializing auth or refreshing tokens.""" + if not profile: + return None + config = configparser.ConfigParser() + config.read(os.getenv("DATABRICKS_CONFIG_FILE", str(pathlib.Path.home() / ".databrickscfg"))) + return config.get(profile, "host", fallback=None) + + +def inventory(source: pathlib.Path, *, profile: str | None, host: str | None) -> dict[str, Any]: + project = AgentProject.load(source) + receipts = load_receipts(source) + resources: list[dict[str, Any]] = [] + + def add(kind: str, name: str | None, action: str | None = None) -> None: + if not name: + return + receipt = next( + ( + r + for r in receipts + if r["host"] == (host or "").rstrip("/") and r["kind"] == kind and r["name"] == name + ), + None, + ) + resources.append( + { + "kind": kind, + "name": name, + "id": receipt["id"] if receipt else None, + "ownership": "created" if receipt and receipt["created"] else "adopted_or_unknown", + "verification": "not_checked", + "url": None, + "next_action": action, + "cleanup": receipt.get("cleanup", "active") if receipt else None, + } + ) + + add( + "deployment", + _prefixed_name(project.deployment_name) if project.deployment_name else None, + "Run agentbricks deploy to create or update the deployment.", + ) + add( + "memory_store", + project.memory_store, + "Run agentbricks deploy to provision the declared store.", + ) + add( + "session_store", + project.session_store, + "Run agentbricks deploy to provision the declared store.", + ) + add( + "experiment", + project.trace_experiment_name, + "Run agentbricks deploy to resolve the tracing experiment.", + ) + for tool in project.tools: + add( + f"tool:{tool.source.kind}", + tool.source.service or tool.source.function or tool.source.space_id or tool.id, + "Tool access is verified when invoked; binding alone does not establish readiness.", + ) + # Include earlier deployment names and stores after the manifest changes, so orphaned resources + # are visible. Receipts from other workspaces must not appear as this workspace's resources. + for receipt in receipts: + if receipt["host"] == (host or "").rstrip("/") and not any( + r["kind"] == receipt["kind"] and r["name"] == receipt["name"] for r in resources + ): + add(receipt["kind"], receipt["name"]) + resources[-1]["previous_binding"] = True + return { + "schema_version": 1, + "project": str(source.resolve()), + "framework": project.framework.value, + "server": project.server.value, + "profile": profile, + "workspace": host, + "verification": "not_checked", + "resources": resources, + "next_actions": ( + ["Select --profile to identify the workspace."] if not profile else [] + ) + + ( + ["Run agentbricks deploy to name and deploy this project."] + if not project.deployment_name + else [] + ) + + [ + "Run agentbricks status --verify for read-only resource checks.", + "Resource availability does not verify model/tool invocation or the deployed app's permissions.", + ], + } + + +def verify_inventory(document: dict[str, Any], client: Any, profile: str) -> None: + for resource in document["resources"]: + kind, name = resource["kind"], resource["name"] + try: + if kind == "deployment": + app = client.workspace_client.apps.get(name) + resource["id"] = app.service_principal_client_id + resource["url"] = app.url or f"{client.host}/apps/{quote(name, safe='')}" + state = getattr(getattr(app, "app_status", None), "state", None) + resource["state"] = getattr(state, "value", state) + elif kind == "memory_store": + store = _resolve_memory_store(client, name) + if store is None: + raise AgentCliError( + "Store was not found in accessible stores.", error_code="NOT_FOUND" + ) + resource["id"] = render.field(store, "name") + elif kind == "session_store": + resource["id"] = render.field(client.get_session_store(name), "name") + elif kind == "experiment": + mlflow = _mlflow() + _set_tracking_uri(mlflow, profile) + experiment = mlflow.get_experiment_by_name(name) + if experiment is None or experiment.lifecycle_stage == "deleted": + raise AgentCliError("Experiment does not exist.", error_code="NOT_FOUND") + resource["id"] = experiment.experiment_id + resource["url"] = experiment_url(client.host, experiment.experiment_id) + else: + continue + resource["verification"] = "accessible" + resource["next_action"] = None + except Exception as exc: # noqa: BLE001 - report partial results without losing the inventory + resource["verification"] = ( + "missing" + if isinstance(exc, NotFound) + or getattr(exc, "error_code", None) in {"NOT_FOUND", "RESOURCE_DOES_NOT_EXIST"} + else "error" + ) + resource["error"] = str(exc) + if resource["verification"] == "error": + resource["next_action"] = ( + "Resolve the reported authentication/permission/service error and retry status --verify." + ) + document["verification"] = "checked" + document["next_actions"] = [a for a in document["next_actions"] if "status --verify" not in a] + + +def _show(document: dict[str, Any], output: str) -> None: + if output == "json": + render.emit_json(document) + return + click.echo(f"Project: {document['project']}") + click.echo(f"Profile: {document['profile'] or 'not selected'}") + click.echo(f"Workspace: {document['workspace'] or 'unresolved'}") + for resource in document["resources"]: + click.echo( + f" {resource['kind']}: {resource['name']} [{resource['verification']}; {resource['ownership']}]" + ) + for key in ("id", "url", "state", "error", "next_action"): + if resource.get(key): + click.echo(f" {key}: {resource[key]}") + for action in document["next_actions"]: + click.echo(f" Next: {action}") + + +@click.command() +@_source_option +@click.option( + "--verify", is_flag=True, help="Check resource availability using read-only workspace APIs." +) +@click.pass_obj +def status(obj, source: pathlib.Path, verify: bool) -> None: + """Show project bindings and resources; configuration alone is not verified readiness.""" + client = None + if verify: + if not obj.profile: + raise AgentCliError("Select --profile before verifying workspace resources.") + client = obj.client() + host = client.host if client is not None else _configured_host(obj.profile) + document = inventory(source, profile=obj.profile, host=host) + if client is not None: + verify_inventory(document, client, obj.profile) + _show(document, obj.output) + + +def cleanup_plan(document: dict[str, Any], receipts: list[dict[str, Any]]) -> list[dict[str, Any]]: + plan = [] + for resource in document["resources"]: + receipt = next( + ( + r + for r in receipts + if r["host"] == document["workspace"].rstrip("/") + and r["kind"] == resource["kind"] + and r["name"] == resource["name"] + ), + None, + ) + removable = bool(receipt and receipt["created"] and resource["kind"] == "deployment") + plan.append( + { + "kind": resource["kind"], + "name": resource["name"], + "action": "delete" if removable else "retain", + "reason": "Created by this project; app identity must still match before deletion." + if removable + else "Shared-capable resource or no creation receipt; retained to protect other consumers.", + "runtime_store": "delete owner-validated managed store" + if removable and receipt and receipt.get("managed_runtime") + else "retain", + "result": receipt.get("cleanup", "active") if receipt else "retained", + } + ) + return plan + + +def apply_cleanup( + source: pathlib.Path, plan: list[dict[str, Any]], receipts: list[dict[str, Any]], client: Any +) -> bool: + failed = False + for item in plan: + if item["action"] != "delete" or item["result"] == "deleted": + continue + receipt = next( + r + for r in receipts + if r["host"] == client.host.rstrip("/") + and r["kind"] == item["kind"] + and r["name"] == item["name"] + ) + name = receipt["name"] + try: + try: + app = client.workspace_client.apps.get(name) + except NotFound: + # Without the app, there is no current owner to validate. Never delete a residual + # Runtime Store just because its name matches the receipt. + item["result"] = "app_missing" + item["reason"] = ( + "App already absent; any remaining Runtime Store is retained for manual inspection." + ) + else: + if app.service_principal_client_id != receipt["id"]: + raise AgentCliError( + "App identity changed since creation; retaining the app and Runtime Store." + ) + if receipt.get("managed_runtime"): + lakebase_runtime_store.delete(client, name, receipt["id"]) + client.workspace_client.apps.delete(name) + item["result"] = "deleted" + receipt["cleanup"] = item["result"] + receipt.pop("cleanup_error", None) + except Exception as exc: # noqa: BLE001 - keep independent resources and retries usable + failed = True + item["result"] = "failed" + item["error"] = str(exc) + receipt["cleanup"] = "failed" + receipt["cleanup_error"] = str(exc) + save_receipts(source, receipts) + return failed + + +@click.command() +@_source_option +@click.option("--apply", is_flag=True, help="Apply the previewed cleanup; default is preview only.") +@click.option("--yes", "-y", is_flag=True, help="Skip the confirmation prompt (requires --apply).") +@click.pass_obj +def cleanup(obj, source: pathlib.Path, apply: bool, yes: bool) -> None: + """Preview cleanup of project-created deployments; retain shared stores, traces and tools.""" + if yes and not apply: + raise AgentCliError("--yes requires --apply; cleanup is preview-only by default.") + if not obj.profile: + raise AgentCliError("Select --profile to scope cleanup to one workspace.") + client = obj.client() if apply else None + host = client.host if client is not None else _configured_host(obj.profile) + if not host: + raise AgentCliError("Could not resolve the selected profile's workspace host.") + document = inventory(source, profile=obj.profile, host=host) + receipts = load_receipts(source) + plan = cleanup_plan(document, receipts) + if apply: + # Always show the exact destructive scope before confirmation, including when JSON is the + # requested final output. stderr keeps stdout machine-readable. + targets = [p["name"] for p in plan if p["action"] == "delete" and p["result"] != "deleted"] + click.echo( + f"Workspace: {host}\nDelete: {', '.join(targets) or '(none)'}\nShared data, tools, source files and local state are retained.", + err=True, + ) + if targets: + _confirm_destroy( + "Delete these project-created deployments and their managed Runtime Stores", + assume_yes=yes, + err=True, + ) + failed = apply_cleanup(source, plan, receipts, client) + else: + failed = False + result = {"workspace": host, "profile": obj.profile, "applied": apply, "resources": plan} + if obj.output == "json": + render.emit_json(result) + else: + click.echo(f"Cleanup {'results' if apply else 'preview'} — {host}") + for item in plan: + click.echo( + f" {item['action']}: {item['kind']} {item['name']} — {item['result']}. {item['reason']}" + ) + if item.get("error"): + click.echo(f" {item['error']}") + if not apply: + click.echo( + "Run cleanup --apply to review and confirm deletion. No resources were changed." + ) + if failed: + raise click.exceptions.Exit(1) diff --git a/integrations/agentbricks/src/databricks_agentbricks/evaluation.py b/integrations/agentbricks/src/databricks_agentbricks/evaluation.py new file mode 100644 index 000000000..a93eaf3be --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/evaluation.py @@ -0,0 +1,219 @@ +"""Small offline-dataset evaluation runner for the managed runtime's actual HTTP agent.""" + +from __future__ import annotations + +import argparse +import json +import os +import pathlib +from typing import Any +from uuid import uuid4 + +import click + +from databricks_agentbricks.cli.endpoint import _authorization_header, _resolve_endpoint +from databricks_agentbricks.cli.endpoint_transport import EndpointRequest, HttpSession +from databricks_agentbricks.errors import AgentCliError + + +def load_cases(path: pathlib.Path) -> list[dict[str, Any]]: + cases = [] + for number, line in enumerate(path.read_text().splitlines(), 1): + if not line.strip(): + continue + try: + case = json.loads(line) + if ( + not isinstance(case.get("inputs", {}).get("query"), str) + or not case["inputs"]["query"].strip() + ): + raise ValueError("inputs.query must be a nonempty string") + expected = case.get("expectations", {}).get("contains", []) + if not isinstance(expected, list) or not all( + isinstance(s, str) and s for s in expected + ): + raise ValueError("expectations.contains must be a list of nonempty strings") + cases.append(case) + except (ValueError, TypeError, AttributeError) as exc: + raise ValueError(f"Invalid evaluation case at {path}:{number}: {exc}") from exc + if not cases: + raise ValueError("The evaluation dataset is empty.") + return cases + + +def response_text(response: Any) -> str: + """Read only final assistant text from the managed adapter response, never tool output.""" + if not isinstance(response, dict) or response.get("status") != "completed": + raise ValueError("Agent invocation did not complete successfully.") + payload = response.get("output") + if not isinstance(payload, dict) or payload.get("status") != "completed": + raise ValueError("Agent did not produce a completed answer (it may require approval).") + texts = [] + for message in payload.get("output", []): + if not isinstance(message, dict) or message.get("role", message.get("type")) not in { + "assistant", + "ai", + }: + continue + content = message.get("content", "") + if isinstance(content, str): + texts.append(content) + elif isinstance(content, list): + texts.extend( + block["text"] + for block in content + if isinstance(block, dict) + and block.get("type") in {"text", "output_text"} + and isinstance(block.get("text"), str) + ) + return "\n".join(text for text in texts if text).strip() + + +def invoke_case( + url: str, query: str, *, authorization: str | None, timeout: float +) -> dict[str, Any]: + session_id = str(uuid4()) + headers = {"Content-Type": "application/json", "X-Routing-Key": session_id} + if authorization: + headers["Authorization"] = authorization + try: + response = HttpSession().send( + EndpointRequest( + url=f"{url}/api/invocations", + method="POST", + headers=headers, + timeout=timeout, + body={ + "id": str(uuid4()), + "session_id": session_id, + "input": {"messages": [{"role": "user", "content": query}]}, + "stream": False, + "background": False, + }, + body_set=True, + ) + ) + if response.status_code != 200: + # Do not log an HTML auth page, response headers, or arbitrary server error bodies. + raise ValueError(f"Agent returned HTTP {response.status_code}; inspect the agent logs.") + return { + "answer": response_text(response.body), + "error": None, + "latency_seconds": response.elapsed_seconds, + } + except (AgentCliError, ValueError) as exc: + return {"answer": "", "error": str(exc), "latency_seconds": None} + + +def run_evaluation( + cases: list[dict[str, Any]], + *, + url: str, + authorization: str | None, + timeout: float, + tracking_uri: str, + experiment: str, +) -> tuple[str, bool]: + import mlflow + from mlflow.genai.scorers import scorer + + mlflow.set_tracking_uri(tracking_uri) + mlflow.set_experiment(experiment) + outcomes: list[dict[str, Any]] = [] + + @mlflow.trace(name="evaluate_project_agent", span_type="AGENT") + def predict_fn(query: str) -> dict[str, Any]: + output = invoke_case(url, query, authorization=authorization, timeout=timeout) + outcomes.append(output) + return output + + @scorer + def completed_answer(outputs: dict[str, Any]) -> bool: + return not outputs.get("error") and bool(outputs.get("answer", "").strip()) + + @scorer + def expected_content(outputs: dict[str, Any], expectations: dict[str, Any]) -> bool: + return ( + not outputs.get("error") + and bool(outputs.get("answer", "").strip()) + and all( + text.casefold() in outputs["answer"].casefold() + for text in expectations.get("contains", []) + ) + ) + + # genai.evaluate logs each input/output/error and scorer result, plus aggregate metrics. The + # HTTP route invokes the project's real agent, including its tools and instructions. + # MLflow's default validation invokes the first case an extra time. Skip that preflight: the + # agent may call real tools, so an unscored duplicate would create unnecessary side effects. + key = "MLFLOW_GENAI_EVAL_SKIP_TRACE_VALIDATION" + previous = os.environ.get(key) + os.environ[key] = "true" + try: + result = mlflow.genai.evaluate( + data=cases, + predict_fn=predict_fn, + scorers=[completed_answer, expected_content], + ) + finally: + if previous is None: + os.environ.pop(key, None) + else: + os.environ[key] = previous + passed = len(outcomes) == len(cases) and all( + not row["error"] and row["answer"] for row in outcomes + ) + passed = passed and all( + result.metrics.get(f"{name}/mean", 0) == 1 + for name in ("completed_answer", "expected_content") + ) + return result.run_id, passed + + +def main(argv: list[str] | None = None) -> int: + parser = argparse.ArgumentParser( + description="Evaluate the project agent against a JSONL dataset." + ) + parser.add_argument("--data", type=pathlib.Path, default=pathlib.Path("evals/cases.jsonl")) + target = parser.add_mutually_exclusive_group() + target.add_argument("--url", help="Local agent URL (default: http://127.0.0.1:8000).") + target.add_argument("--app", help="Databricks deployment name; requires --profile with OAuth.") + parser.add_argument("--profile", help="Explicit Databricks OAuth profile for --app.") + parser.add_argument("--tracking-uri", default="sqlite:///.agentbricks/evaluations.db") + parser.add_argument("--experiment", default="agentbricks-starter-evaluation") + parser.add_argument("--timeout", type=float, default=120) + args = parser.parse_args(argv) + if args.app and not args.profile: + parser.error("--app requires --profile ") + if args.timeout <= 0: + parser.error("--timeout must be positive") + try: + cases = load_cases(args.data) + url, authenticate = _resolve_endpoint( + args.app, None if args.app else args.url or "http://127.0.0.1:8000", args.profile + ) + authorization = _authorization_header(args.profile) if authenticate else None + if args.tracking_uri == "sqlite:///.agentbricks/evaluations.db": + pathlib.Path(".agentbricks").mkdir(exist_ok=True) + run_id, passed = run_evaluation( + cases, + url=url, + authorization=authorization, + timeout=args.timeout, + tracking_uri=args.tracking_uri, + experiment=args.experiment, + ) + click.echo( + f"MLflow run: {run_id}\nTracking URI: {args.tracking_uri}\nResult: {'PASS' if passed else 'FAIL'}" + ) + if args.tracking_uri.startswith("sqlite:"): + click.echo( + f"View per-case results: uv run mlflow ui --backend-store-uri {args.tracking_uri}" + ) + return 0 if passed else 1 + except (OSError, ValueError, AgentCliError) as exc: + parser.exit(2, f"Evaluation failed: {exc}\n") + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/integrations/agentbricks/src/databricks_agentbricks/project_resources.py b/integrations/agentbricks/src/databricks_agentbricks/project_resources.py new file mode 100644 index 000000000..63e9d48aa --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/project_resources.py @@ -0,0 +1,89 @@ +"""Local provisioning receipts. A binding or a matching name is never proof of ownership.""" + +from __future__ import annotations + +import json +import os +import pathlib +import tempfile +from typing import Any + +from databricks_agentbricks.errors import AgentCliError + +_RECEIPTS = pathlib.Path(".agentbricks/resources.json") +_KINDS = {"deployment", "memory_store", "session_store"} + + +def load_receipts(source: pathlib.Path) -> list[dict[str, Any]]: + path = source / _RECEIPTS + if not path.exists(): + return [] + try: + document = json.loads(path.read_text()) + if document.get("schema_version") != 1 or not isinstance(document.get("resources"), list): + raise ValueError("unsupported receipt schema") + for item in document["resources"]: + if ( + not isinstance(item, dict) + or item.get("kind") not in _KINDS + or not all( + isinstance(item.get(key), str) and item[key] for key in ("host", "name", "id") + ) + or not isinstance(item.get("created"), bool) + ): + raise ValueError("invalid resource receipt") + return document["resources"] + except (OSError, ValueError, TypeError, AttributeError) as exc: + raise AgentCliError(f"Could not read resource receipts at {path}: {exc}.") from exc + + +def save_receipts(source: pathlib.Path, resources: list[dict[str, Any]]) -> None: + path = source / _RECEIPTS + path.parent.mkdir(parents=True, exist_ok=True) + temporary: str | None = None + try: + with tempfile.NamedTemporaryFile(mode="w", dir=path.parent, delete=False) as output: + temporary = output.name + json.dump({"schema_version": 1, "resources": resources}, output, indent=2) + output.write("\n") + os.replace(temporary, path) + except OSError as exc: + if temporary: + pathlib.Path(temporary).unlink(missing_ok=True) + raise AgentCliError( + f"Could not save provisioning receipts at {path}: {exc}.", + hint="The cloud operation may have succeeded. Retain this project's receipts and inspect status before retrying.", + ) from exc + + +def record_resource( + source: pathlib.Path, + *, + host: str, + kind: str, + name: str, + resource_id: str, + created: bool, + managed_runtime: bool = False, +) -> None: + """Record successful create/reuse immediately, including deployments that later fail.""" + resources = load_receipts(source) + host = host.rstrip("/") + previous = next( + (r for r in resources if (r["host"], r["kind"], r["name"]) == (host, kind, name)), + None, + ) + receipt = { + "host": host, + "kind": kind, + "name": name, + "id": resource_id, + "created": created + or bool(previous and previous["id"] == resource_id and previous["created"]), + "managed_runtime": managed_runtime, + "cleanup": "active", + } + if previous is not None: + resources.remove(previous) + resources.append(receipt) + save_receipts(source, resources) diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/README.md b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/README.md index 8d28c444f..667c0a91c 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/README.md +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/README.md @@ -142,3 +142,12 @@ replacement attempt fails with `MCP_USER_AUTH_RECOVERY_UNSUPPORTED` before agent interruptions remain unsupported and fail with `MCP_USER_AUTH_HITL_UNSUPPORTED`. The agent's existing namespaced memory, conversation store, and checkpointer behavior is unchanged; OBO does not add another saver. + +## Project status, cleanup and evaluations + +Use `agentbricks --profile status` for a non-mutating inventory; add `--verify` for +read-only workspace checks. `agentbricks --profile cleanup` previews safe cleanup, +and `cleanup --apply` asks before deleting project-created deployments. Shared data is retained. + +Run `uv run python evals/run.py` against the local agent to record a starter evaluation in MLflow. +See [evals/README.md](evals/README.md) for per-case results, additional cases and deployed runs. diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/README.md b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/README.md new file mode 100644 index 000000000..7be6d3f83 --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/README.md @@ -0,0 +1,62 @@ +# Evaluate the project agent + +Start this project in another terminal with `agentbricks --profile dev`. +Then run from the project directory: + +```bash +uv run python evals/run.py +``` + +This sends each case to **your running project agent** at `/api/invocations`: its instructions, +model and tools run normally. The dataset is offline (checked into this project); model/tool +calls still use the configured workspace and can incur normal charges. Each case gets a fresh +session so cases cannot inherit conversation history. The script never overrides the agent model. + +The starter supports the OpenAI Agents SDK and LangGraph **Agent Bricks server** templates. +Custom HTTP server templates need an evaluator matching their own input/output contract. + +Results, per-case input/output/errors, scorer feedback and aggregate pass rates are saved in a +local MLflow experiment. The script prints the run ID and UI command: + +```bash +uv run mlflow ui --backend-store-uri sqlite:///.agentbricks/evaluations.db +``` + +Open http://127.0.0.1:5000, select `agentbricks-starter-evaluation`, then the printed run. +A failed invocation, empty answer, approval interruption or unmet expectation causes a nonzero exit +status. Authentication/HTTP failures remain visible as failed cases; they are not counted as passes. +These simple deterministic smoke checks are a starting point, not a measure of domain quality. + +## Extend the dataset + +Add one JSON object per line to `evals/cases.jsonl`. `inputs.query` is sent to the agent. +`expectations.contains` is a list of case-insensitive substrings required in the final assistant +answer. For example: + +```json +{"inputs":{"query":"What is 3 + 4? Reply with the number only."},"expectations":{"contains":["7"]}} +``` + +Change the prompts and expectations for your agent's real use cases. Custom scorers, tool-quality +checks and LLM judges can be added to `run_evaluation` or a project-specific copy; see +[MLflow GenAI evaluation](https://mlflow.org/docs/latest/genai/eval-monitor/). + +## Evaluate a deployment or share results + +```bash +uv run python evals/run.py --app --profile +uv run python evals/run.py --url http://127.0.0.1:8000 --data evals/cases.jsonl --timeout 180 +``` + +`--app` uses the selected OAuth profile for Databricks Apps; `--url` is unauthenticated and intended +for local testing. To log evaluations in your chosen workspace rather than locally: + +```bash +uv run python evals/run.py --app --profile \ + --tracking-uri databricks:// --experiment /Users//agent-evaluations +``` + +Keep the local `.agentbricks` directory to retain your evaluation history. Project cleanup retains +local results, workspace experiments, session/memory stores and tools because other projects may +use them. The evaluator invokes the actual agent, so use isolated test resources for tool cases +that can write data or perform external actions. diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/cases.jsonl b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/cases.jsonl new file mode 100644 index 000000000..69804af2d --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-langgraph/evals/cases.jsonl @@ -0,0 +1,2 @@ +{"inputs":{"query":"What is 2 + 2? Reply with the number only."},"expectations":{"contains":["4"]}} +{"inputs":{"query":"Reply with exactly: evaluation-ready"},"expectations":{"contains":["evaluation-ready"]}} diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/README.md b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/README.md index 3dfd9a268..1a98acfd0 100644 --- a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/README.md +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/README.md @@ -144,3 +144,12 @@ replacement attempt fails with `MCP_USER_AUTH_RECOVERY_UNSUPPORTED` before agent interruptions remain unsupported and fail with `MCP_USER_AUTH_HITL_UNSUPPORTED`, without storing a credential-bearing `RunState`. Existing namespaced memory and conversation-store behavior is unchanged; OBO does not add another saver. + +## Project status, cleanup and evaluations + +Use `agentbricks --profile status` for a non-mutating inventory; add `--verify` for +read-only workspace checks. `agentbricks --profile cleanup` previews safe cleanup, +and `cleanup --apply` asks before deleting project-created deployments. Shared data is retained. + +Run `uv run python evals/run.py` against the local agent to record a starter evaluation in MLflow. +See [evals/README.md](evals/README.md) for per-case results, additional cases and deployed runs. diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/README.md b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/README.md new file mode 100644 index 000000000..7be6d3f83 --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/README.md @@ -0,0 +1,62 @@ +# Evaluate the project agent + +Start this project in another terminal with `agentbricks --profile dev`. +Then run from the project directory: + +```bash +uv run python evals/run.py +``` + +This sends each case to **your running project agent** at `/api/invocations`: its instructions, +model and tools run normally. The dataset is offline (checked into this project); model/tool +calls still use the configured workspace and can incur normal charges. Each case gets a fresh +session so cases cannot inherit conversation history. The script never overrides the agent model. + +The starter supports the OpenAI Agents SDK and LangGraph **Agent Bricks server** templates. +Custom HTTP server templates need an evaluator matching their own input/output contract. + +Results, per-case input/output/errors, scorer feedback and aggregate pass rates are saved in a +local MLflow experiment. The script prints the run ID and UI command: + +```bash +uv run mlflow ui --backend-store-uri sqlite:///.agentbricks/evaluations.db +``` + +Open http://127.0.0.1:5000, select `agentbricks-starter-evaluation`, then the printed run. +A failed invocation, empty answer, approval interruption or unmet expectation causes a nonzero exit +status. Authentication/HTTP failures remain visible as failed cases; they are not counted as passes. +These simple deterministic smoke checks are a starting point, not a measure of domain quality. + +## Extend the dataset + +Add one JSON object per line to `evals/cases.jsonl`. `inputs.query` is sent to the agent. +`expectations.contains` is a list of case-insensitive substrings required in the final assistant +answer. For example: + +```json +{"inputs":{"query":"What is 3 + 4? Reply with the number only."},"expectations":{"contains":["7"]}} +``` + +Change the prompts and expectations for your agent's real use cases. Custom scorers, tool-quality +checks and LLM judges can be added to `run_evaluation` or a project-specific copy; see +[MLflow GenAI evaluation](https://mlflow.org/docs/latest/genai/eval-monitor/). + +## Evaluate a deployment or share results + +```bash +uv run python evals/run.py --app --profile +uv run python evals/run.py --url http://127.0.0.1:8000 --data evals/cases.jsonl --timeout 180 +``` + +`--app` uses the selected OAuth profile for Databricks Apps; `--url` is unauthenticated and intended +for local testing. To log evaluations in your chosen workspace rather than locally: + +```bash +uv run python evals/run.py --app --profile \ + --tracking-uri databricks:// --experiment /Users//agent-evaluations +``` + +Keep the local `.agentbricks` directory to retain your evaluation history. Project cleanup retains +local results, workspace experiments, session/memory stores and tools because other projects may +use them. The evaluator invokes the actual agent, so use isolated test resources for tool cases +that can write data or perform external actions. diff --git a/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/cases.jsonl b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/cases.jsonl new file mode 100644 index 000000000..69804af2d --- /dev/null +++ b/integrations/agentbricks/src/databricks_agentbricks/templates/agent-openai/evals/cases.jsonl @@ -0,0 +1,2 @@ +{"inputs":{"query":"What is 2 + 2? Reply with the number only."},"expectations":{"contains":["4"]}} +{"inputs":{"query":"Reply with exactly: evaluation-ready"},"expectations":{"contains":["evaluation-ready"]}} diff --git a/integrations/agentbricks/tests/e2e/evaluation_smoke.py b/integrations/agentbricks/tests/e2e/evaluation_smoke.py new file mode 100644 index 000000000..0b8ac7799 --- /dev/null +++ b/integrations/agentbricks/tests/e2e/evaluation_smoke.py @@ -0,0 +1,107 @@ +"""Real HTTP runtime + local MLflow smoke test, without model or cloud calls. + +Run with full MLflow installed (either framework extra): + uv run python tests/e2e/evaluation_smoke.py --output /tmp/agentbricks-eval-smoke + +This verifies the evaluator's transport, case failure handling, extension and MLflow persistence. +Use the scaffold's evals/run.py against a real project agent separately to verify model/tool calls. +""" + +from __future__ import annotations + +import argparse +import json +import pathlib +import socket +import threading +import time + +import mlflow +import uvicorn + +from databricks_agentbricks.evaluation import load_cases, run_evaluation +from databricks_agentkit import DurableAgentServer + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--output", type=pathlib.Path, required=True) + args = parser.parse_args() + args.output.mkdir(parents=True, exist_ok=True) + calls = [] + app = DurableAgentServer() + + @app.invoke + async def invoke(value, context): + query = value["messages"][-1]["content"] + calls.append({"query": query, "session_id": context.session_id}) + return {"status": "completed", "output": [{"role": "assistant", "content": query}]} + + with socket.socket() as sock: + sock.bind(("127.0.0.1", 0)) + port = sock.getsockname()[1] + server = uvicorn.Server(uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning")) + worker = threading.Thread(target=server.run, daemon=True) + worker.start() + deadline = time.monotonic() + 15 + while not server.started and worker.is_alive() and time.monotonic() < deadline: + time.sleep(0.01) + if not server.started: + raise RuntimeError("Runtime did not start") + tracking_uri = f"sqlite:///{(args.output / 'evaluations.db').resolve()}" + results = [] + try: + for label, expected_pass, cases in ( + ( + "failing-expectation", + False, + [{"inputs": {"query": "hello"}, "expectations": {"contains": ["missing"]}}], + ), + ( + "passing-and-extended", + True, + [ + {"inputs": {"query": query}, "expectations": {"contains": [query]}} + for query in ("hello", "evaluation-ready", "additional-case") + ], + ), + ): + path = args.output / f"{label}.jsonl" + path.write_text("\n".join(json.dumps(case) for case in cases) + "\n") + run_id, passed = run_evaluation( + load_cases(path), + url=f"http://127.0.0.1:{port}", + authorization=None, + timeout=10, + tracking_uri=tracking_uri, + experiment="evaluator-smoke", + ) + assert passed == expected_pass, (label, passed) + run = mlflow.get_run(run_id) + mlflow.flush_trace_async_logging() + traces = mlflow.search_traces( + locations=[run.info.experiment_id], run_id=run_id, return_type="list" + ) + assert len(traces) == len(cases) + assert all( + trace.data.spans + and {"completed_answer", "expected_content"} + <= {assessment.name for assessment in trace.info.assessments} + for trace in traces + ) + results.append( + {"scenario": label, "run_id": run_id, "passed": passed, "metrics": run.data.metrics} + ) + assert len({call["session_id"] for call in calls}) == len(calls) + assert len(calls) == 4, "Each case should invoke the agent exactly once." + finally: + server.should_exit = True + worker.join(timeout=10) + (args.output / "results.json").write_text( + json.dumps({"tracking_uri": tracking_uri, "runs": results, "calls": calls}, indent=2) + ) + print(json.dumps(results, indent=2)) # noqa: T201 - standalone e2e report + + +if __name__ == "__main__": + main() diff --git a/integrations/agentbricks/tests/unit_tests/evaluation_test.py b/integrations/agentbricks/tests/unit_tests/evaluation_test.py new file mode 100644 index 000000000..707d5c52d --- /dev/null +++ b/integrations/agentbricks/tests/unit_tests/evaluation_test.py @@ -0,0 +1,118 @@ +"""Evaluation validates datasets and grades final assistant answers, not transport success.""" + +import json + +import pytest + +from databricks_agentbricks import evaluation +from databricks_agentbricks.cli.endpoint_transport import EndpointResponse +from databricks_agentbricks.cli.init import _copy_packaged_template + + +def answer(value="hello", *, status="completed"): + return { + "status": "completed", + "output": { + "status": status, + "output": [ + {"role": "tool", "content": "tool result must not pass a case"}, + {"role": "assistant", "content": value}, + ], + }, + } + + +def test_extracts_only_final_assistant_text(): + assert evaluation.response_text(answer()) == "hello" + assert ( + evaluation.response_text( + answer( + [ + {"type": "text", "text": "visible"}, + {"type": "reasoning", "text": "opaque"}, + {"type": "output_text", "text": "answer"}, + ] + ) + ) + == "visible\nanswer" + ) + with pytest.raises(ValueError, match="approval"): + evaluation.response_text(answer(status="interrupted")) + with pytest.raises(ValueError, match="successfully"): + evaluation.response_text({"status": "failed", "error": "boom"}) + + +def test_langgraph_wire_format_extracts_ai_messages_and_excludes_tools(): + response = { + "status": "completed", + "output": { + "status": "completed", + "output": [ + {"type": "human", "content": "What is 2 + 2?"}, + {"type": "tool", "content": "incorrect tool result"}, + {"type": "ai", "content": "4"}, + ], + }, + } + assert evaluation.response_text(response) == "4" + + +def test_real_request_contract_uses_unique_sessions_and_keeps_project_model(monkeypatch): + requests = [] + + def send(self, request): + requests.append(request) + return EndpointResponse(request.url, 200, {}, answer(), 0.1) + + monkeypatch.setattr(evaluation.HttpSession, "send", send) + for _ in range(2): + assert ( + evaluation.invoke_case("http://localhost:8000", "hello", authorization=None, timeout=5)[ + "answer" + ] + == "hello" + ) + assert requests[0].body["session_id"] != requests[1].body["session_id"] + assert requests[0].body["input"] == {"messages": [{"role": "user", "content": "hello"}]} + assert requests[0].headers["X-Routing-Key"] == requests[0].body["session_id"] + + +def test_http_error_record_does_not_log_sensitive_response_body(monkeypatch): + monkeypatch.setattr( + evaluation.HttpSession, + "send", + lambda _, r: EndpointResponse(r.url, 401, {}, "credential-containing-auth-page", 0.1), + ) + output = evaluation.invoke_case( + "http://localhost:8000", "hello", authorization="Bearer secret", timeout=5 + ) + assert "HTTP 401" in output["error"] + assert "secret" not in str(output) + assert "credential" not in str(output) + + +@pytest.mark.parametrize( + "case", + [{}, {"inputs": {"query": " "}}, {"inputs": {"query": "x"}, "expectations": {"contains": "x"}}], +) +def test_invalid_dataset_fails_before_invocation(tmp_path, case): + path = tmp_path / "cases.jsonl" + path.write_text(json.dumps(case)) + with pytest.raises(ValueError, match="Invalid evaluation case"): + evaluation.load_cases(path) + + +def test_app_requires_explicit_profile(): + with pytest.raises(SystemExit) as error: + evaluation.main(["--app", "test"]) + assert error.value.code == 2 + + +@pytest.mark.parametrize("framework", ["openai", "langgraph"]) +def test_generated_evaluator_does_not_require_unreleased_runtime_module(tmp_path, framework): + destination = tmp_path / framework + _copy_packaged_template(f"agent-{framework}", destination) + source = (destination / "evals/run.py").read_text() + assert "from databricks_agentbricks.evaluation" not in source + assert "def run_evaluation(" in source + assert len(evaluation.load_cases(destination / "evals/cases.jsonl")) == 2 diff --git a/integrations/agentbricks/tests/unit_tests/project_lifecycle_test.py b/integrations/agentbricks/tests/unit_tests/project_lifecycle_test.py new file mode 100644 index 000000000..9c1e67732 --- /dev/null +++ b/integrations/agentbricks/tests/unit_tests/project_lifecycle_test.py @@ -0,0 +1,259 @@ +"""Status never provisions; cleanup requires creation evidence and current remote ownership.""" + +import json +from types import SimpleNamespace +from unittest.mock import Mock + +import pytest +from click.testing import CliRunner +from databricks.sdk.errors import NotFound + +from databricks_agentbricks.agent_project import AgentProject +from databricks_agentbricks.cli import deploy as deploy_commands +from databricks_agentbricks.cli import project as commands +from databricks_agentbricks.errors import AgentCliError +from databricks_agentbricks.project_resources import load_receipts, record_resource + + +@pytest.fixture +def project(tmp_path): + project = AgentProject.create( + tmp_path, + framework="openai", + server="agentbricks", + memory_store="memory", + session_store="sessions", + experiment_name="/Shared/test", + ) + project.set_deployment_name("demo") + project.write() + return tmp_path + + +@pytest.fixture +def context(monkeypatch): + monkeypatch.setattr(commands, "_configured_host", lambda profile: "https://test.example") + client = Mock(host="https://test.example") + client.workspace_client.apps.get.return_value = SimpleNamespace( + service_principal_client_id="principal-1", url="https://demo.example", app_status=None + ) + return SimpleNamespace(profile="test", output="json", client=Mock(return_value=client)) + + +def receipt(project, **kwargs): + record_resource( + project, + host="https://test.example", + kind="deployment", + name="agent-bricks-demo", + resource_id="principal-1", + created=True, + managed_runtime=True, + **kwargs, + ) + + +def test_status_and_cleanup_preview_never_initialize_auth_or_mutate(project, context): + receipt(project) + before = {p: p.read_bytes() for p in project.rglob("*") if p.is_file()} + runner = CliRunner() + status = runner.invoke(commands.status, ["--source", str(project)], obj=context) + assert status.exit_code == 0, status.output + document = json.loads(status.output) + assert document["verification"] == "not_checked" + assert all(r["verification"] == "not_checked" for r in document["resources"]) + preview = runner.invoke(commands.cleanup, ["--source", str(project)], obj=context) + assert preview.exit_code == 0, preview.output + plan = json.loads(preview.output) + assert [r["kind"] for r in plan["resources"] if r["action"] == "delete"] == ["deployment"] + assert context.client.call_count == 0 + assert before == {p: p.read_bytes() for p in project.rglob("*") if p.is_file()} + + +def test_verify_retains_partial_results_and_permission_error(project, context, monkeypatch): + client = context.client() + client.get_session_store.side_effect = AgentCliError( + "No access", error_code="PERMISSION_DENIED" + ) + monkeypatch.setattr(commands, "_resolve_memory_store", lambda *_: {"name": "memory-stores/123"}) + monkeypatch.setattr(commands, "_set_tracking_uri", lambda *_: None) + monkeypatch.setattr( + commands, "_mlflow", lambda: SimpleNamespace(get_experiment_by_name=lambda _: None) + ) + before = (project / "agent.toml").read_bytes() + result = CliRunner().invoke( + commands.status, ["--source", str(project), "--verify"], obj=context + ) + assert result.exit_code == 0, result.output + by_kind = {r["kind"]: r for r in json.loads(result.output)["resources"]} + assert by_kind["deployment"]["verification"] == "accessible" + assert by_kind["memory_store"]["id"] == "memory-stores/123" + assert by_kind["session_store"]["verification"] == "error" + assert by_kind["experiment"]["verification"] == "missing" + client.create_session_store.assert_not_called() + client.create_memory_store.assert_not_called() + assert before == (project / "agent.toml").read_bytes() + + +def test_cancelled_cleanup_deletes_nothing(project, context): + receipt(project) + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply"], obj=context, input="n\n" + ) + assert result.exit_code == 1 + context.client().workspace_client.apps.delete.assert_not_called() + context.client().delete_runtime_store.assert_not_called() + assert load_receipts(project)[0]["cleanup"] == "active" + + +def test_adopted_and_other_workspace_apps_are_never_deleted(project, context): + record_resource( + project, + host="https://test.example", + kind="deployment", + name="agent-bricks-demo", + resource_id="principal-1", + created=False, + ) + record_resource( + project, + host="https://other.example", + kind="deployment", + name="agent-bricks-other", + resource_id="principal-other", + created=True, + ) + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code == 0, result.output + assert all(r["action"] == "retain" for r in json.loads(result.stdout)["resources"]) + context.client().workspace_client.apps.delete.assert_not_called() + + +def test_recreated_app_with_same_name_is_retained(project, context): + receipt(project) + context.client().workspace_client.apps.get.return_value.service_principal_client_id = ( + "replacement" + ) + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code == 1 + assert "identity changed" in result.stdout + context.client().workspace_client.apps.delete.assert_not_called() + context.client().delete_runtime_store.assert_not_called() + + +def test_partial_failure_retry_removes_runtime_before_app_and_is_idempotent( + project, context, monkeypatch +): + receipt(project) + client = context.client() + events = [] + + def delete_runtime(*_): + events.append("runtime") + + monkeypatch.setattr(commands.lakebase_runtime_store, "delete", delete_runtime) + client.workspace_client.apps.delete.side_effect = [RuntimeError("try again"), None] + runner = CliRunner() + args = ["--source", str(project), "--apply", "--yes"] + first = runner.invoke(commands.cleanup, args, obj=context) + assert first.exit_code == 1 + assert load_receipts(project)[0]["cleanup"] == "failed" + second = runner.invoke(commands.cleanup, args, obj=context) + assert second.exit_code == 0, second.output + assert load_receipts(project)[0]["cleanup"] == "deleted" + third = runner.invoke(commands.cleanup, args, obj=context) + assert third.exit_code == 0, third.output + assert client.workspace_client.apps.delete.call_count == 2 + assert events == ["runtime", "runtime"] + + +def test_runtime_failure_retains_app(project, context, monkeypatch): + receipt(project) + monkeypatch.setattr( + commands.lakebase_runtime_store, "delete", Mock(side_effect=AgentCliError("Wrong owner")) + ) + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code == 1 + context.client().workspace_client.apps.delete.assert_not_called() + + +def test_missing_app_never_deletes_unvalidated_runtime(project, context): + receipt(project) + context.client().workspace_client.apps.get.side_effect = NotFound("gone") + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code == 0, result.output + assert "app_missing" in result.stdout + context.client().delete_runtime_store.assert_not_called() + + +def test_receipt_reuse_preserves_created_ownership_but_replacement_does_not(project): + receipt(project) + record_resource( + project, + host="https://test.example", + kind="deployment", + name="agent-bricks-demo", + resource_id="principal-1", + created=False, + ) + assert load_receipts(project)[0]["created"] is True + record_resource( + project, + host="https://test.example", + kind="deployment", + name="agent-bricks-demo", + resource_id="principal-2", + created=False, + ) + assert load_receipts(project)[0]["created"] is False + + +@pytest.mark.parametrize( + "payload", ["[]", "null", '{"schema_version":2}', '{"schema_version":1,"resources":[{}]}'] +) +def test_corrupt_receipts_fail_closed(project, context, payload): + directory = project / ".agentbricks" + directory.mkdir() + (directory / "resources.json").write_text(payload) + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code != 0 + context.client().workspace_client.apps.delete.assert_not_called() + + +def test_store_receipt_survives_later_provisioning_failure(project, context, monkeypatch): + monkeypatch.setattr( + deploy_commands, "_ensure_memory_store", lambda *_: ({"name": "memory-stores/123"}, True) + ) + monkeypatch.setattr( + deploy_commands, + "_ensure_session_store", + Mock(side_effect=AgentCliError("Service unavailable")), + ) + with pytest.raises(AgentCliError, match="Service unavailable"): + deploy_commands._reconcile_declared_stores( + "memory", "sessions", context.client(), source=project + ) + stored = load_receipts(project) + assert len(stored) == 1 + assert stored[0]["kind"] == "memory_store" and stored[0]["created"] + + +def test_apply_uses_authenticated_workspace_instead_of_configured_host(project, context): + receipt(project) + context.client().host = "https://other.example" + result = CliRunner().invoke( + commands.cleanup, ["--source", str(project), "--apply", "--yes"], obj=context + ) + assert result.exit_code == 0, result.output + assert all(r["action"] == "retain" for r in json.loads(result.stdout)["resources"]) + context.client().workspace_client.apps.delete.assert_not_called()