diff --git a/integrations/agentbricks/CONTRIBUTING.md b/integrations/agentbricks/CONTRIBUTING.md index 3853ab120..9063f74d7 100644 --- a/integrations/agentbricks/CONTRIBUTING.md +++ b/integrations/agentbricks/CONTRIBUTING.md @@ -101,6 +101,32 @@ the same change, in `src/databricks_agentbricks/cli/doctor.py`: Also refresh the framework references and examples in `cli.md` and `README.md`, and add doctor test coverage for the new framework in `tests/unit_tests/doctor_test.py`. +## Live tool tests + +Read-only tool discovery can be checked against the installed wheel without creating a +project or deploying an agent: + +```sh +AGENTBRICKS_E2E_PROFILE= .venv-functional/bin/pytest tests/e2e/tool_discovery_test.py -v +``` + +The checks compare default and MCP-filtered discovery with the compatibility service +list. Set `AGENTBRICKS_E2E_SCHEMA=catalog.schema` to exercise another schema. Local +add/review/remove flows and help pages are covered by `tests/functional/cli_smoke_test.py`. + +The opt-in Genie tests exercise both frameworks against an existing Genie space. From +`integrations/agentbricks`, with both framework extras installed: + +```sh +DATABRICKS_CONFIG_PROFILE=my-workspace RUN_AGENTBRICKS_GENIE_TESTS=1 \ + AGENTBRICKS_GENIE_SPACE_ID=SPACE_ID \ + uv run pytest tests/integration_tests/genie_tools_test.py +``` + +By default they ask for the row count of `samples.nyctaxi.trips`. Set +`AGENTBRICKS_GENIE_QUESTION` for another dataset and +`AGENTBRICKS_GENIE_EXPECTED_VALUE` to assert a known result cell. + ## Cutting a release Run **Cut Agent Bricks release** from the Actions tab with a version such as `0.4.0` or diff --git a/integrations/agentbricks/README.md b/integrations/agentbricks/README.md index ae8ab835d..af9bd5a2c 100644 --- a/integrations/agentbricks/README.md +++ b/integrations/agentbricks/README.md @@ -1,75 +1,22 @@ # Agent Bricks CLI (`agentbricks`) Agent Bricks CLI is an experimental command-line interface for building and deploying custom -agents on Databricks. It manages memory, sessions, tracing, and deployments from one authenticated -command. +agents on Databricks. It manages memory, sessions, tracing, and deployments through one CLI. > The underlying APIs are in preview and may need workspace enablement. -## Overview - -A managed path from your custom agent code to a production-ready, scalable, durable agent hosted on -Databricks in minutes - with no server framework to build, no infrastructure to provision, and no -invocation protocol to design yourself. Bring your own agent, or start from a template. - -- **Deployment** - a guided lifecycle (scaffold, run locally, deploy) that turns an agent project - into a hosted endpoint. Databricks provisions the compute, the stores your agent binds (session, - memory), and the access grants, so you ship application code and get a running endpoint. -- **Runtime** - a managed HTTP invocation contract (synchronous, streaming, background) plus optional - durable execution (persistence, heartbeats, crash recovery) backed by Databricks Lakebase, with no - database or job queue to operate. Use the opinionated `DurableAgentServer` to get it out of the box, - or bring your own server for full control. `AgentApp` remains available as a deprecated - compatibility alias; new code should use `DurableAgentServer`. - -**Deployment** - -![Deployment: from a blank directory to a running service](docs/deployment.svg) - -- **Agent project** - `agentbricks init` scaffolds a deployable project from a framework template - (LangGraph or OpenAI Agents) with the runtime, tests, and an optional chat UI wired up; you edit - the application code (model, tools, prompts). -- **`agent.toml`** - the declarative source of truth for the Databricks-managed infrastructure your - agent depends on: tool bindings (data sandbox, managed MCP services, Unity Catalog functions) and - memory, session, and durability resources. `agentbricks deploy` reads it to provision and wire everything - up (detailed under [Agent tools](#agent-tools)). -- **`agentbricks deploy`** - provisions the bound stores, grants the app's service principal access to - them, provisions the durable-runtime database when durability is on, configures tracing, and rolls - out the app. `agentbricks deployments` covers the lifecycle (list, get, logs, start, stop, delete). -- **`agentbricks dev`** - runs your agent from the same manifest the deployment uses, so local behavior - matches what ships. - -**Runtime** - -![Runtime: one FastAPI server, run as DurableAgentServer or your own implementation](docs/runtime.svg) - -The two ways to run an agent: - -- **`DurableAgentServer` - opinionated, batteries included.** Register one handler and get the managed - invocation contract (synchronous, streaming, background). Enable the durable runtime so - long-running and background work survives restarts, redeploys, and crashes. The framework - templates are thin layers over `DurableAgentServer` (HTTP contract detailed under [Runtime](#runtime)). -- **Custom server - generic, full control.** `agentbricks init --server custom` scaffolds a minimal FastAPI - server with no `DurableAgentServer`: you define your own endpoints, request/response shapes, and protocol. - `agentbricks dev` and `agentbricks deploy` run and ship it the same way. - ## Prerequisites -- **Python ≥3.10** — the `agentbricks` CLI installs and runs on any Python 3.10+. The - `memory`, `sessions`, `tools`, and `agentbricks tracing bind`/`unbind` commands need nothing else. -- **[`uv`](https://docs.astral.sh/uv/)** — needed to scaffold, run, and deploy an - agent (`agentbricks init` → `agentbricks dev` → `agentbricks deploy`): the scaffolded project builds - its environment and launches with `uv run`, both locally and in the deployed Apps - runtime. The store/session/tools commands and `agentbricks tracing bind`/`unbind` don't need it; - `agentbricks tracing list`/`get` do, to read `agentbricks dev`'s local trace store. -- **[Databricks CLI](https://docs.databricks.com/dev-tools/cli/)** — needed for - browser-based `agentbricks login`. If a profile is already authenticated, the CLI uses it - directly and the Databricks CLI is optional. +- **Python ≥3.10** to run the CLI; generated agent projects require Python 3.11+. +- **[`uv`](https://docs.astral.sh/uv/)** to scaffold, run, and deploy an agent. Resource + commands can run without it; reading local traces requires it. +- **[Databricks CLI](https://docs.databricks.com/dev-tools/cli/)** for browser-based + `agentbricks login`. An already-authenticated profile does not require it. ## Installation -From PyPI: - The `databricks-agentbricks` Python distribution installs the `agentbricks` command and AgentKit. +Install it from PyPI: ```sh pip install databricks-agentbricks @@ -84,910 +31,131 @@ pip install 'git+https://github.com/databricks/databricks-ai-bridge.git#subdirec The base package includes the CLI, store SDK, and `DurableAgentServer` HTTP runtime. Generated projects declare their framework dependencies automatically. -## Shell completion -Add this to `~/.zshrc`: -```sh -eval "$(_AGENTBRICKS_COMPLETE=zsh_source agentbricks)" -``` - -## Authentication - -Agent Bricks CLI uses [Databricks authentication](https://docs.databricks.com/aws/en/dev-tools/cli/authentication). -Ask the CLI to authenticate and remember a named profile: - -```sh -agentbricks login --profile -agentbricks sessions stores list -``` - -`agentbricks login` validates existing credentials first. If credentials are missing or rejected in -an interactive terminal, the CLI runs `databricks auth login --profile `, revalidates the -profile, and stores the selection in the existing `~/.agentbricks/config.json` state file. This -browser-based setup requires the Databricks CLI. In non-interactive environments, authenticate the -profile before running `agentbricks`. `agentbricks logout` forgets the saved selection without revoking the underlying -credentials. - -If Databricks SDK default authentication is already configured, you can skip `agentbricks login`. -You can also pass the global `--profile/-p` option before an individual command, for example -`agentbricks --profile tools list`. Use `--output json` for scripting. - ## Quickstart -The shortest path from a blank directory to a running and deployed agent: +Create a new agent project with LangGraph (the default framework). Use `--framework openai` +instead for an **OpenAI Agents SDK** project. This chooses the agent framework, not the model +provider; both templates call a Databricks AI Gateway model. ```sh -agentbricks login --profile -agentbricks init my-agent -cd my-agent -agentbricks dev # run locally -agentbricks deploy my-agent # deploy to Databricks -``` - -`agentbricks dev` runs the agent locally on `http://localhost:8000`, wrapping the Databricks Apps -local runtime so local behavior matches a deployment. - -`agentbricks deploy my-agent` deploys a Databricks App named `agent-bricks-my-agent`, provisions the -stores declared in `agent.toml`, and grants the app's service principal access to the stores and -direct App-auth tool resources declared there. Use `agentbricks deployments list` to find deployed -apps, and `agentbricks deployments get agent-bricks-my-agent` to print an app's URL and status. - -`agentbricks init` declares default memory and session stores in `agent.toml`, so the deployed agent has -long-term memory and durable conversation history. It creates `-<6-letter-token>-memory` and -`-<6-letter-token>-sessions`, and records both names in `agent.toml`. -`agentbricks deploy` creates them if they don't exist yet. Point the agent at stores you already have with -`agentbricks memory bind ` / `agentbricks sessions bind `, or scaffold without stores using -`agentbricks init --server custom` (see [Initialize the chat app demo](#initialize-the-chat-app-demo)). - -To exercise the agent (locally under `agentbricks dev` or once deployed), `agentbricks endpoint invoke` sends -it an HTTP request. MLflow tracing is on by default (`agentbricks init` binds a default -`/Shared/agentbricks_traces/` experiment): `agentbricks dev` traces to a local MLflow server under -`.agentbricks/` and `agentbricks deploy` to the bound workspace experiment; `agentbricks tracing list` shows the -available traces. - -## Public names - -Use `agentbricks` for the CLI and `AgentKitClient` from `databricks_agentkit` for the Python SDK. New -projects store local state under `.agentbricks/` and `~/.agentbricks/`, use -`server = "agentbricks"` in `agent.toml`, and deploy apps with the `agent-bricks-` prefix. - -## AgentKit SDK - -`AgentKitClient` adds a small resource-oriented layer over AgentKit APIs. Pass it an -authenticated Databricks `WorkspaceClient`, or omit the argument to use the -Databricks SDK's default authentication resolution: - -```python -from databricks.sdk import WorkspaceClient -from databricks_agentkit import AgentKitClient - -agentkit = AgentKitClient(WorkspaceClient(profile="my-workspace")) - -session_store = agentkit.session_stores.create("support-agent-sessions") -session = session_store.add(actor_id="customer-123", session_id="case-456") -session.append_items( - [ - {"type": "message", "role": "user", "content": "I need help with my cluster."}, - {"type": "message", "role": "assistant", "content": "Let's take a look."}, - ] -) - -memory_store = agentkit.memory_stores.create("coding-agent-memory") -memory = memory_store.add( - actor_id="alice", - path="/preferences/style.md", - content="The user prefers concise answers.", -) -results = memory_store.search( - actor_id="alice", - query="response preferences", - limit=10, -) -memory = memory.update(content="The user prefers very concise answers.") -memory.delete() -``` - -The root collections manage stores: `agentkit.memory_stores.create/get/list` and -`agentkit.session_stores.create/get/list`. A returned store owns operations on its -contents, such as `memory_store.add()`, `memory_store.get("memory-id")`, -`memory_store.list()`, and `memory_store.search()`, or `session_store.add()`, -`session_store.get("session-id")`, and `session_store.list()`. Returned memories, -sessions, and stores own their `update()` and `delete()` operations. - -All `list()` methods return iterators that automatically consume server pages. List -`page_size` and search `limit` values must be between 1 and 100. `session.list_items()` -also auto-pages. `session.fork(...)` creates an independent copy, optionally through -a specific item. Deleting a session with descendants requires -`session.delete(force=True)` to cascade the deletion. - -The resource layer intentionally does not mirror every API method. Its private -transport will be replaced by the generated `WorkspaceClient.mason` service when that -is released, without changing this public surface. Deployment, sandbox, tracing, and -the existing CLI commands remain separate. - -## Runtime - -`DurableAgentServer` runs your agent through one HTTP API for synchronous, streaming, and background -invocations. Register an `@app.invoke` handler, publish progress with `await context.emit(event)`, -and return a JSON result. You can also add your own FastAPI endpoints. - -Start from a template, edit the agent code in `agent/`, and run it locally before deploying: - -```sh -agentbricks init my-agent --framework langgraph --server agentbricks --profile +agentbricks init my-agent --framework langgraph --profile cd my-agent +agentbricks login --profile agentbricks dev -# Stop the local server when ready to deploy. -agentbricks --profile deploy my-agent -``` - -Use `--framework openai` for OpenAI Agents. Templates keep agent code separate from the runtime -adapter and declare default Session and Memory Store bindings in `agent.toml`. - -Each managed run is an **invocation**. Send a client-generated UUID `id` and your agent's `input`: - -```json -{ - "id": "550e8400-e29b-41d4-a716-446655440000", - "session_id": "support-case-123", - "input": {"messages": [{"role": "user", "content": "Hello"}]}, - "background": true, - "stream": true -} -``` - -The optional top-level `session_id` groups invocations into one application session. It is distinct -from the invocation `id` and from the `X-Routing-Key` sticky-routing header. - -| Endpoint | Behavior | -| --- | --- | -| `POST /api/invocations` | Defaults to synchronous execution: `200` with the result under `output`. `stream: true` returns SSE events. `background: true` returns `202` with a status URL; adding `stream: true` also includes an events URL. | -| `GET /api/invocations/{id}` | Returns the invocation status and, when completed, its output. | -| `GET /api/invocations/{id}/events?after={cursor}` | Streams events after the last received event ID, allowing clients to reconnect. | - -The UUID also acts as an idempotency key: repeating the same request reuses the existing invocation -while its record is retained; using the ID for a different request returns `409`. - -`agentbricks dev` keeps execution state in process and loses it on restart. For projects with -`[agent].server = "agentbricks"`, `agentbricks deploy` provisions a persistent Runtime Store for requests, -status, events, and results. Register `@app.recover` to restart interrupted app-auth work after -worker failures. Recovery is at-least-once, so external side effects must be idempotent. Session -and Memory Stores separately preserve the state used by your agent. - -The managed path uses the internal Runtime Store API to create a dedicated database in the -workspace's shared Lakebase project and give the app SP ownership. The managed runtime initializes its schema and -tables; no manual Lakebase grant or Postgres app-resource attachment is needed. Backend selection -is an internal rollout detail, not a user-facing setting; the current implementation retains the legacy -per-app Lakebase project by default. Once enabled, redeploy reads the stored backend and verifies -the app identity, and `agentbricks deployments delete` removes the managed store before deleting the app. -The switch does not migrate existing deployments between backends. Managed cleanup errors retain -the app for retry. Direct app deletion bypasses managed store cleanup. - -For a tool using `auth = "user"`, the Runtime Store still records token-free invocation state, -events, and results. The forwarded user credential remains process-local for the active attempt and -is never written to the Runtime Store. A replacement attempt after failure recovery stops with -`MCP_USER_AUTH_RECOVERY_UNSUPPORTED` because the original request credential is no longer present. - -Use `server = "custom"` to deploy your own HTTP server without provisioning a Runtime Store. -Changing the server type of an existing deployment is not supported. To use a different server, -scaffold a new project with the desired `agentbricks init --server` option and deploy it under a new name. -See the [runtime guide](src/databricks_agentkit/runtime/README.md) for agent hooks, full API examples, -and recovery behavior. - -## Memory and sessions - -To hold context, an agent needs two kinds of state: the state of the interaction it is handling right -now, and the durable knowledge it carries from one conversation to the next. Databricks provides a -fully managed store for each, both backed by Lakebase and usable from agents built on any framework: - -- **Managed agent sessions** store an agent's session state: the state an agent or framework keeps - for one interaction. Most commonly this is the conversation history (the ordered transcript of - messages, tool calls, and results), but it can be any state a framework persists, such as a - LangGraph graph. The agent reads it at the start of a turn and appends to it as the interaction - runs. -- **Managed agent memory** stores durable facts, preferences, and decisions that an agent recalls in - later, separate conversations, retrieved by semantic search. - -The examples below use the [`AgentKitClient` Python SDK](#agentkit-sdk); the same operations are available -as `agentbricks sessions` / `agentbricks memory` CLI commands. - -![Sessions and memory: the agent reads and appends one conversation's transcript in the session store, and recalls and saves durable facts in the memory store, which outlive any single conversation.](docs/sessions_and_memory.png) - -### Sessions - -A **session store** holds **sessions**, and each session holds an ordered list of **session items**. A -session is one interaction — typically a conversation thread — grouped under an `actor_id` (who it -belongs to; set this from trusted application context, never a model- or user-supplied value) and -identified by a caller-chosen `session_id` (the service generates one if you omit it). Each item is an -opaque, JSON-compatible `data` value — a message, tool call, result, or reasoning block — that -Databricks stores and returns verbatim, in order, and never mutates once appended. - -Create a store, start a session, append the conversation's turns, and read the history back on a later -request: - -```python -from databricks.sdk import WorkspaceClient -from databricks_agentkit import AgentKitClient - -agentkit = AgentKitClient(WorkspaceClient()) - -session_store = agentkit.session_stores.create("support-agent-sessions") -session = session_store.add(actor_id="customer-123", session_id="case-456") - -session.append_items( - [ - {"type": "message", "role": "user", "content": "I need help with my cluster."}, - {"type": "message", "role": "assistant", "content": "Let's take a look."}, - ] -) - -# On a later turn, reload the session and read its full history in order. -session = session_store.get("case-456") -history = [item.data for item in session.list_items()] # list_items auto-pages -``` - -A session can be **forked** into an independent branch: a new session seeded with the original's -history, linked back to its origin by `parent_session_id`. Fork the full history, or only up to a -specific item, to explore an alternate continuation without disturbing the original thread: - -```python -branch = session.fork(actor_id="customer-123") # add up_to_item_id=... to branch up to one item -``` - -Deleting a session that has such descendants requires `session.delete(force=True)` to cascade. - -In an agent configured with `server = "agentbricks"`, you don't call these directly — the framework adapter reads and appends -session state for you. With LangGraph, pass `checkpointer()` when you build the agent and scope each -run with `thread_config(session_id)`; the OpenAI Agents adapter exposes the same as -`session_store(session_id)`: - -```python -from databricks_agentkit.langgraph import checkpointer, thread_config - -agent = create_agent(model=..., tools=[...], checkpointer=checkpointer()) -result = await agent.ainvoke(inputs, config=thread_config(session_id)) -``` - -### Memory - -A **memory store** holds **memory entries**. Each entry is a free-form `content` string plus a short -`description` used for retrieval, keyed by three fields: `actor_id` (whose memory it is — set from -trusted application context, never a model- or user-supplied value), `path` (a filesystem-like key -within an actor, such as `/preferences/response-style.md`), and an optional `session_id` (the session -an entry came from, for provenance). An entry is uniquely identified by its `actor_id`, `path`, and -optional `session_id`. - -Write an entry when the agent learns something durable, then recall it in a later, separate -conversation with a natural-language search — results are ranked by full-text (BM25) relevance, up to -100 entries, with no pagination or vector similarity: - -```python -from databricks.sdk import WorkspaceClient -from databricks_agentkit import AgentKitClient - -agentkit = AgentKitClient(WorkspaceClient()) - -memory_store = agentkit.memory_stores.create("support-agent-memory") -memory_store.add( - actor_id="user-123", - path="/preferences/communication.md", - content="Prefers email over phone. Timezone: PST.", - description="User 123 communication preferences", -) - -# In a later, separate conversation, recall what the agent knows about this user. -results = memory_store.search(actor_id="user-123", query="communication preferences", limit=10) -``` - -To browse rather than search, `memory_store.list(actor_id=..., path_prefix=...)` returns entries -directly. - -In an agent configured with `server = "agentbricks"`, add the memory tools so the model can read and write memory during a run. -`memory_tools(actor)` exposes `remember` and `recall` bound to one actor's partition; it resolves the -store from the `[memory_store]` binding, carried to the runtime by the `AGENT_MEMORY_STORE` env var -that `agentbricks deploy` injects, and returns no tools when no store is set, so the agent runs unchanged. -That "no store set" path is also how it runs under `agentbricks dev`, which runs locally: memory is off -there (the store is provisioned and used only at deploy). The OpenAI Agents adapter exposes the same as -`memory_tools()`: - -```python -from databricks_agentkit.langgraph import memory_tools - -agent = create_agent(model=..., tools=[*your_tools, *memory_tools(actor)]) ``` -> **`actor_id` partitions data; it is not access control.** Both stores are workspace-scoped and -> authorized at the store level, so any principal that can reach a store can read and write every -> actor's entries. For strict isolation between tenants or users, use a separate store per boundary. -> Grant another principal — such as your app's service principal — access with -> `session_store.grant_permission(principal_id)` or `memory_store.grant_permission(principal_id)`; -> `agentbricks deploy` does this for the deployed app automatically. - -### Declaring and provisioning stores - -For a deployed agent, `agent.toml` declares which stores it uses and `agentbricks deploy` provisions them — -you don't create stores by hand. `agentbricks init` declares a default memory and session store named from -the project; override those names, point at stores you already have, or let `deploy` create them: - -```sh -# Scaffold a project with default memory and session stores declared in agent.toml. -agentbricks init my-agent - -# Override the declared store names at init time. -agentbricks init my-agent --memory-store support-agent-memory --session-store support-agent-sessions - -# Or point an existing project at specific stores (edits agent.toml only; creates nothing). -agentbricks sessions bind support-agent-sessions -agentbricks memory bind support-agent-memory - -# deploy creates any declared-but-missing store and grants the app's service principal access. -agentbricks deploy my-agent -``` - -Memory and session stores are independent resources: deleting one never affects the other. - -## Commands - -For the full command reference - every command, subcommand, argument, and option, in table form - -see [`cli.md`](cli.md). The tree below is a quick overview. - -```text -agentbricks [-p ] [-o text|json] - login [--profile P] - logout - init [--framework openai|langgraph] [--server agentbricks|custom] - [--disable-chat-app] - [--memory-store NAME] [--session-store NAME] - [--existing] [--profile P] [directory] - doctor [directory] - dev [--source PATH] [--prepare-environment] [--app-port PORT] - memory - bind STORE [--source PATH] - unbind [--source PATH] - stores create | list | get | update | delete - entries create | get | list | search | update | delete - sessions create | list | get | update | delete | fork - bind STORE [--source PATH] - unbind [--source PATH] - stores create | list | get | update | delete - items list | append | pop | clear - tracing - bind (--experiment-name NAME | --experiment-id ID) [--source PATH] - unbind [--source PATH] - list | get [--experiment-name NAME | --experiment-id ID] [--source PATH] - tools - add sandbox --scope SCOPE [--scope SCOPE ...] - [--no-databricks-access-token-included] [--source PATH] - add mcp SERVICE [--name NAME] [--source PATH] - add uc-function FUNCTION [--name NAME] [--source PATH] - add genie-one [--name NAME] [--auth user|app] [--source PATH] - add genie-agent SPACE_ID [--name NAME] [--auth user|app] [--source PATH] - list [--kind sandbox|mcp|uc-function|genie-one|genie-agent] - [--schema CATALOG.SCHEMA] - remove TOOL_ID [MCP_SERVICE] [--source PATH] - deploy [] [--source PATH] [--instances N] - deployments list | get | logs | start | stop | delete - endpoint - invoke [APP] --path PATH [--url URL] [--json JSON] [--sse] -``` - -## Bring an existing agent - -From the existing project, prepare a migration for your coding agent: +`init` copies a project with a chat UI and records the profile in its local `.env`. `login` +authenticates that profile for the CLI. Open the URL printed by `dev` (usually +`http://localhost:8000`) and send a message. Stop `dev` with Ctrl-C, then deploy: ```sh -agentbricks init --framework langgraph --existing . -agentbricks init --framework openai --existing . +agentbricks deploy my-agent --source . +agentbricks deployments get agent-bricks-my-agent ``` -Before or after the conversion, inspect its progress without changing the repository or contacting -Databricks: - -```sh -agentbricks doctor . -agentbricks -o json doctor . -``` - -Doctor exits 0 only when the project has a valid Agent Bricks manifest and matching project -metadata, uses the Agent Bricks server, declares the framework-appropriate `databricks-agentbricks` -extra and a non-empty `app.yaml` command, constructs `DurableAgentServer` with an `invoke` hook, and -calls a recognized adapter for the selected framework in production Python source. Test, example, and -old/stale directories do not count as source evidence. A failed report is the normal result for a -project that still needs migration; run -`agentbricks init --framework --existing ` with the appropriate framework -to prepare the migration instructions. Doctor never imports or executes the target's source, and a -bounded source scan that exceeds a limit is reported while the evidence it already found still counts. -Its findings are static repository evidence, not proof that the configured startup command executes -the files it finds. - -This writes `agent-bricks-migrate/` containing a skill, a prompt to paste into your coding agent, -`references/migration.json`, and a reference project generated from the templates bundled with the -installed CLI. The bundle sits outside any single agent's configuration directory; `.claude/skills/` -and `.agent/skills/` each receive a small skill that points at it, so Claude Code, Codex, and -similar tools discover the same instructions without duplicating the reference. Agent Bricks CLI -prepares the instructions; the coding agent performs and verifies the conversion. Init leaves application -source, dependencies, `.env`, and existing `.agentbricks/project.toml` configuration intact and refuses to overwrite -existing migration files. - -The bundle is scaffolding for the migration, not part of the application: delete `agent-bricks-migrate/` -and the two pointer skills once the conversion is done, and keep them out of commits meanwhile. - -The skill follows the shared managed-runtime contract for the selected framework, included in new projects and -migration references: -[LangGraph](src/databricks_agentbricks/templates/agent-langgraph/AGENTKIT_CONTRACT.md) or -[OpenAI Agents SDK](src/databricks_agentbricks/templates/agent-openai/AGENTKIT_CONTRACT.md). It explicitly -handles existing history, custom state and output, recovery, and client/session contracts. For -LangGraph, switching checkpointers does not migrate old conversations (likewise, the OpenAI Agents -SDK keeps prior Session transcripts and RunState behind); unresolved transitions require a user -decision. - -The reference honors `--disable-chat-app`, `--memory-store`, `--session-store`, and the selected -profile. These are migration intent; init does not provision resources or change the existing -application. Migration supports LangGraph and the OpenAI Agents SDK with the managed server (`server = "agentbricks"`); -`--server custom` is not supported for `--existing`. - -## Invoke HTTP endpoints - -`agentbricks endpoint invoke` is a low-level HTTP command. It resolves and authenticates a deployed -Databricks App, or targets localhost and arbitrary servers through `--url`. It does not assume an -agent protocol: provide the method, path, query parameters, and complete JSON body required by the -server. - -```sh -agentbricks --profile endpoint invoke agent-bricks-my-agent \ - --path /api/invocations \ - --json '{"id":"00000000-0000-4000-8000-000000000001","session_id":"support-case-123","input":[{"role":"user","content":"Hello"}]}' +`deploy` creates or updates the Databricks App and provisions the stores declared in +`agent.toml`. It attempts to grant the App access to them; inspect its output for access or +tracing warnings. `deployments get` prints the App URL and status. Open the deployed chat UI +and send a message to verify it. -agentbricks endpoint invoke --url http://localhost:8000 \ - --path /api/invocations \ - --json '{"id":"00000000-0000-4000-8000-000000000001","session_id":"support-case-123","input":[{"role":"user","content":"Hello"}]}' -``` +### Invoke from the command line -The JSON body remains explicit even for generated agents. For example, managed runtime agents require -a client-generated invocation ID, and streaming servers require their own streaming field plus -`--sse` so the CLI consumes the response as Server-Sent Events. +The chat UI is one client of the generated agent. To call its API directly, send a new UUID as +`id` for each turn. Both generated frameworks require a nonempty top-level `session_id` on every +request; reuse it for turns in the same conversation: ```sh -SESSION_ID=$(uuidgen) -INVOCATION_ID=$(uuidgen) -agentbricks --profile endpoint invoke agent-bricks-my-agent \ - --path /api/invocations \ - --routing-key "$SESSION_ID" \ - --json "{\"id\":\"$INVOCATION_ID\",\"session_id\":\"$SESSION_ID\",\"input\":[{\"role\":\"user\",\"content\":\"Run the report\"}]}" - -INVOCATION_ID=$(uuidgen) -agentbricks --profile endpoint invoke agent-bricks-my-agent \ +SESSION_ID=$(python3 -c 'import uuid; print(uuid.uuid4())') +INVOCATION_ID=$(python3 -c 'import uuid; print(uuid.uuid4())') +agentbricks endpoint invoke agent-bricks-my-agent \ --path /api/invocations \ - --routing-key "$SESSION_ID" \ - --sse \ - --json "{\"id\":\"$INVOCATION_ID\",\"session_id\":\"$SESSION_ID\",\"input\":[{\"role\":\"user\",\"content\":\"Hello\"}],\"stream\":true}" -``` - -`--routing-key` keeps a session on one app replica (sticky routing): set it to your stable session id -and it is sent verbatim in the `X-Routing-Key` request header. It is routing only and is never used -as the session id - put the session id at the top level of the `--json` body for session continuity. This -also works with a direct App URL and with the generated runtime on localhost. OAuth and session headers are managed by -the runtime; arbitrary custom request headers are intentionally not exposed by this command. - -## Command help - -Use the conventional help flag at any command level. Every command's help includes runnable -examples: - -```sh -agentbricks --help -agentbricks deploy --help -agentbricks sessions items append --help -``` - -## Agent tools - -For projects with `[agent].server = "agentbricks"` (the default from `agentbricks init`), `agent.toml` is the -declarative source of truth for Databricks-managed infrastructure: the Runtime Store, sandbox, -managed MCP, Genie and Unity Catalog function bindings, plus memory and session resources. `agentbricks tools -add` updates only this file; direct TOML edits have the same behavior. Both managed-server framework -adapters read the managed bindings at runtime without generating or patching agent source: - -```sh -agentbricks tools add sandbox --scope table:samples.nyctaxi.trips -agentbricks tools add mcp system.ai.web_search -agentbricks tools add uc-function catalog.schema.lookup_ticket -agentbricks tools add genie-one -agentbricks tools add genie-agent SPACE_ID -agentbricks tools remove mcp system.ai.web_search -agentbricks tools list + --json "{\"id\":\"$INVOCATION_ID\",\"session_id\":\"$SESSION_ID\",\"input\":{\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}}" ``` -### Managed tool identity and migration - -`agentbricks tools add mcp`, `agentbricks tools add sandbox`, `agentbricks tools add genie-one`, and -`agentbricks tools add genie-agent` write explicit `auth = "user"` by default. -Use `--auth app` for the App service principal instead. This field is on the tool entry, not -inside `source` or `policy`: +Use `agentbricks endpoint invoke --url http://localhost:8000` instead of the App name to +call `dev` while it is running. Invoking a deployed App requires an OAuth-authenticated +profile; PAT profiles cannot access App routes. See the +[runtime guide](src/databricks_agentkit/runtime/README.md) for streaming, background requests, +and recovery. -```toml -[[tools]] -id = "web_search" -auth = "user" -source = { kind = "mcp", service = "system.ai.web_search" } -``` +## First customization: add a custom tool -Direct UC-function bindings remain app/default identity and do not accept `--auth user`. -Managed-tool add commands write the selected identity to `agent.toml`; inspect that manifest to -review configured bindings. Missing legacy auth continues to mean App identity at runtime; it is -never silently upgraded to user identity. - -### Automatic App-identity access on deploy - -`agentbricks deploy` reconciles least-privilege access for resources explicitly declared by App/default -identity tool bindings. It skips every `auth = "user"` binding because those calls use the request -user's permissions instead of the App service principal. - -| Explicit `agent.toml` resource | Automatic App service-principal access | -| --- | --- | -| UC function | Apps `uc_securable`: `FUNCTION` / `EXECUTE` | -| Genie Agent space | Apps `genie_space`: `CAN_RUN` | -| Sandbox table scope | Apps `uc_securable`: `TABLE` / `SELECT` or `MODIFY` | -| Sandbox volume scope | Apps `uc_securable`: `VOLUME` / `READ_VOLUME` or `WRITE_VOLUME` | -| Sandbox Workspace path | Workspace ACL: `CAN_READ` or `CAN_EDIT` | -| External MCP service | Unity Catalog: effective `EXECUTE` plus `USE_SCHEMA` and `USE_CATALOG` on its named parents | -| Built-in `system.ai` MCP service, including Sandbox and Genie One | Platform-managed access defaults; Agent Bricks does not mutate system securables | - -Native Genie One has no resource identifier in its binding, so it does not add a resource-specific -grant. Use a Genie Agent binding when the App identity should be scoped to one explicit Genie Space. - -Apps-backed tool resources are named deterministically and reconciled to the manifest on each -deploy: removing a binding removes that Agent Bricks-owned Apps resource while preserving Runtime -Store, tracing, and user-owned resources. MCP and Workspace ACL grants are additive in this release -because their permission APIs do not expose trustworthy Agent Bricks ownership metadata; removing -those bindings does not revoke an independently valid grant. - -Only direct resources are automatic. Agent Bricks never discovers or grants tables and warehouses -used by a Genie Space, objects called by a UC function, or resources wrapped by an MCP service. -Grant those transitive dependencies manually when the called service uses the App identity. If any -required direct grant cannot be read, applied, or verified, deploy stops before source upload and -leaves the currently deployed version untouched. - -`DurableAgentServer` derives its request-auth policy directly from the managed tool bindings in -`agent.toml`. Projects do not maintain a separate request-auth contract marker: a managed tool with -`auth = "user"` requires a transient request-user credential. Code-first tools can declare any -additional API scopes that Agent Bricks cannot infer from Python: - -```toml -[auth.user] -required = true -additional_api_scopes = ["sql"] -``` +Add `agent/tools/count_words.py` to the generated project. Use the version for your chosen +framework: -`additional_api_scopes` is additive: deploy unions it with scopes inferred from managed bindings, -deduplicates the result, and preserves unrelated scopes already configured on the App. Scope names -are not restricted to a client-side allowlist; Databricks Apps validates whether a requested scope -is supported. Entries must be non-empty strings without surrounding whitespace or control -characters, and a non-empty list requires `required = true`. This request-auth contract is supported -only with `[agent].server = "agentbricks"`. - -Generated framework adapters pass the request-bound resolver to agent construction. A code-first -tool should obtain its user client from that resolver inside the active invocation rather than -creating or persisting a user credential: +LangGraph: ```python from langchain_core.tools import tool -def sql_tools(workspace_client_for): - @tool - def run_statement(statement: str) -> str: - client = workspace_client_for("user") - response = client.statement_execution.execute_statement( - warehouse_id="...", - statement=statement, - ) - return str(response.result) - - return [run_statement] -``` - -The resolver is request-bound and closes after the attempt. Agent Bricks does not inject it into -arbitrary auto-discovered decorated tools; build those tools from the resolver passed to the -generated request-aware agent function. - -Request-user invocations use the same synchronous, streaming, background, status, event-replay, -and idempotency APIs as app-auth invocations. The Runtime Store records only token-free request -state, events, and results. The forwarded credential stays process-local for the active first -attempt and closes when that attempt completes, fails, or is cancelled. A replacement attempt after -failure recovery stops with `MCP_USER_AUTH_RECOVERY_UNSUPPORTED` because no user credential is -available; neither the invoke nor recovery handler runs for that attempt. - -Before deploying user-auth tools from an older project, migrate its request handler and framework -adapter to the current request-auth-aware `DurableAgentServer` template, then explicitly choose `user` or -`app` on **every** managed MCP, sandbox, or Genie entry. Changing `agent.toml` alone does not -upgrade copied Python adapter code. Outdated adapters fail closed rather than silently using App -identity. App-only legacy projects and generic bring-your-own source directories keep the existing -path. - -Deploy derives Apps user scopes from explicit `auth = "user"` bindings and unions them with -`[auth.user].additional_api_scopes`: - -| Binding | Requested Apps scopes | -| --- | --- | -| Managed MCP (governed ingress) | `ai-gateway` | -| `system.ai.dbsql` | `ai-gateway`, `sql` | -| `system.ai.genie_one_mcp` | `ai-gateway`, `genie` | -| Sandbox with a Volume downscope | `ai-gateway`, `files` | -| First-class Genie One or Genie Agent | `genie` | -| Sandbox with token injection | `ai-gateway`, `workspace.workspace` | -| Sandbox with token injection disabled | `ai-gateway` | - -For example, bind Genie tools in a current project with `server = "agentbricks"`: - -```sh -agentbricks tools add mcp system.ai.genie_one_mcp --auth user -agentbricks tools add genie-agent SPACE_ID --auth user -``` - -Mixed bindings and explicit additions request the union. App-auth and legacy bindings add no user -scopes. These are -explicit service-consent scopes, not a claim that gateway access alone authorizes the downstream -resource. OAuth consent does not grant Unity Catalog privileges: the user still needs access to -the configured Genie Space and its underlying data. - -For a new App, deploy explicitly enables user-token forwarding and includes these scopes in the -initial typed SDK create request before uploading source. An existing App that is missing a required -scope needs one-time explicit permission: - -```sh -agentbricks --profile my-workspace deploy my-agent --allow-user-scope-update +@tool +def count_words(text: str) -> int: + """Count whitespace-separated words in text.""" + return len(text.split()) ``` -The `system.ai.dbsql` managed MCP additionally requests the Apps `sql` user scope. This is full SQL -API consent, not `sql:restricted-query`; read-only enforcement remains the service policy plus the -requesting user's Unity Catalog grants. DBSQL does not use Databricks Connect. - -When a user-auth sandbox has `databricks_access_token_included = true`, it requests the Apps -`workspace.workspace` user scope so the injected credential can call workspace APIs. A sandbox -binding with a Volume downscope additionally requests the Apps `files` user scope. -OAuth consent does not grant Volume access: the requesting user still needs the corresponding -Unity Catalog privileges, and the sandbox downscope remains authoritative. A sandbox binding with -token injection disabled does not request `workspace.workspace`; its other resource-derived scopes -still apply. Databricks Apps rejects the legacy bare `workspace` scope, so Agent Bricks requests -`workspace.workspace`. These scopes are requested only for `auth = "user"`; `auth = "app"` uses -the App service principal's permissions instead. - -Review the target App's scopes and coordinate with its other owners before allowing the update. Once -those scopes are present, later deploys do not need the flag. The CLI preserves unrelated scopes, -updates only user scopes and any explicitly requested instance counts, and checks requested **and -effective** scopes before source rollout. It checks for scope changes since preflight, but Apps -read/write is **not atomic**; this is not a lock or a compare-and-swap guarantee. Polling is bounded -and a mismatch stops source deployment. -Users may need to sign out and **re-consent** after changing scopes; effective-scope verification -does not refresh an existing user's consent. - -Apps may report `iam.access-control:read` and `iam.current-user:read` as implicit effective -scopes. The CLI permits these platform defaults during verification but does not request them -as configurable scopes. Explicitly disabled user-token forwarding stops deployment; enable -forwarding and restart the App compute before retrying. - -Removing a tool or switching back to app-only auth **does not remove Apps scopes**. Remove -unneeded scopes explicitly in Databricks Apps, and verify both configured and effective scopes -before declaring removal complete. The CLI does not send empty-list scope updates: the SDK's -`App.as_dict()` omits empty lists, so that would not prove removal succeeded. No scopes are -managed for generic bring-your-own apps without this managed user contract. - -App-auth tools execute with workload privileges. Restrict App `CAN USE` to callers trusted -for **all** App-auth tools, or deploy those tools separately. Models, custom MCP servers, -Memory/Session Stores, and tracing keep their existing credentials. - -For MCP services, the remove command accepts the same service name as the add command. You can also -remove any binding by its `id` in `agent.toml`, for example `agentbricks tools remove web_search`. -Every successful add (including an already-configured no-op) points you to the target project's -`agent.toml` to review configured managed tools and MCP bindings. With `--source`, the message -points to that project's file. JSON add output includes its path in `manifest`. - -`agentbricks tools list` discovers **available integrations to add**, not configured bindings. By default -it shows built-in add recipes and caller-visible MCP Services in `system.ai`. A recipe may still -need your resources: sandbox scopes, a concrete UC function name, or a Genie Space ID. Genie One -needs no additional argument. `system.ai.sandbox` is represented by its scoped recipe rather than -a second unscoped add command. The list does not enumerate every workspace schema, individual -operations inside MCP services, or custom Python tools. - -`agentbricks tools add mcp` looks up the service in the selected workspace before writing `agent.toml`. -Use `agentbricks --profile tools add mcp ` to select a workspace. A missing service or -failed lookup (including authentication or permission errors) leaves the project unchanged. This -checks service metadata access, not whether every tool can be executed at runtime. Removing local -bindings does not require workspace access. - -```sh -agentbricks tools list -agentbricks tools list --kind mcp -agentbricks tools list --kind mcp --schema main.tools -agentbricks tools list --kind sandbox -agentbricks tools list --kind genie-one -agentbricks tools list --kind genie-agent -agentbricks --output json tools list -``` - -No agent project is required for discovery. MCP discovery uses your Databricks profile; the -`sandbox`, `uc-function`, `genie-one`, and `genie-agent` kind filters show local recipes without -authentication. `--schema` requires `--kind mcp` and replaces the default `system.ai` scope. An -API/authentication failure returns nonzero and marks discovery incomplete, while retaining local -recipes; it is not reported as an empty successful discovery. Listing metadata does not verify -runtime execution permissions. - -**Migration:** the former configured `tools list` view and its `--source` option are removed. -Read `agent.toml` (its `[[tools]]` entries) to inspect configured bindings. Discovery JSON uses -`schema_version: 2`, with `available_tools` (`name`, `kind`, `add_command`), `mcp_schema` (null for -local-only recipes), `complete`, and `errors`. Replace old scripts that read configured-list JSON -with TOML inspection. Replace `agentbricks mcp list [--schema catalog.schema]` with -`agentbricks tools list --kind mcp [--schema catalog.schema]`; the former command is removed. Use -`agentbricks tools list --help` for the new discovery contract. +OpenAI Agents SDK: -Read-only live discovery can be checked against the installed wheel without creating a project -or deploying an agent: +```python +from agents import function_tool -```sh -AGENTBRICKS_E2E_PROFILE= .venv-functional/bin/pytest tests/e2e/tool_discovery_test.py -v -``` -The live checks compare default and MCP-filtered discovery with the compatibility service list. -Set `AGENTBRICKS_E2E_SCHEMA=catalog.schema` to exercise an additional schema. The installed CLI's local -add/review/remove flows and all updated help pages are covered by `tests/functional/cli_smoke_test.py`. - -In managed-server templates, custom Python tools are code-first. Write them with the framework's native -decorator in `agent/tools/`: LangGraph uses `@tool`, while OpenAI Agents uses `@function_tool`. The -templates auto-discover decorated tools from that package and add them to the agent; there is no CLI -command or `agent.toml` entry to keep in sync. Customer-managed MCP servers are likewise ordinary -code in `agent/mcps.py` and are joined with the managed bindings by `mcp_tools(...)` or -`mcp_servers(...)`. - -Projects created with `--server custom` do not auto-discover `agent/tools/` or load managed tool -bindings from `agent.toml`, so `agentbricks tools add` rejects those projects. Wire framework-native Python -tools and MCP servers directly in `agent/agent.py` instead. - -If a manifest with `server = "agentbricks"` contains `source = { kind = "python", ... }`, remove that -`[[tools]]` entry; the decorated tool in `agent/tools/` remains active. `agentbricks dev` and `agentbricks deploy` -do not generate or patch Python tool code, and do not alter the manifest's `[[tools]]` bindings. - -Sandbox scopes default to read-only access. Repeat `--scope` to allow more than one resource, use -`volume:` or `workspace:` for those resource types, and use `--permission read_write` only when the -agent needs writes. Every sandbox call carries this fixed downscope in MCP `_meta`, outside the tool -arguments controlled by the model. New sandbox bindings also expose the selected Databricks -credential to sandbox code by default: - -```toml -[[tools]] -id = "sandbox" -auth = "user" -source = { kind = "sandbox", service = "system.ai.sandbox" } -policy = { downscope = [{ resource = "workspace:/Workspace/Shared", permission = "read_only" }], databricks_access_token_included = true } +@function_tool +def count_words(text: str) -> int: + """Count whitespace-separated words in text.""" + return len(text.split()) ``` -With `databricks_access_token_included = true`, the sandbox receives `DATABRICKS_HOST`, a short-lived -`DATABRICKS_TOKEN`, and `DATABRICKS_AUTH_TYPE`, so code such as -`WorkspaceClient().current_user.me()` can call workspace APIs. This policy does not choose the -identity: `auth = "user"` uses the request user's OBO credential, while `auth = "app"` uses the -Databricks App service principal. Use `--no-databricks-access-token-included` when adding a sandbox that -does not need workspace API access. Existing manifests that omit `databricks_access_token_included` -remain disabled until explicitly updated. +Both templates discover decorated tools in `agent/tools/` automatically. Run `agentbricks dev` +again and ask the local chat UI: “Use count_words to count the words in: the quick brown fox.” +The tool returns `4`. Edit `agent/agent.py` to change the model, instructions, or agent logic. +For Databricks-managed tools, see the [Agent tools guide](docs/agent-tools.md). -### Genie tools - -Genie One and Genie Agent support ship with Agent Bricks CLI, but bindings are opt-in, like sandbox tools. -Installing the current `databricks-agentbricks` distribution does not configure a Genie Space ID or enable a Genie binding. Add only the -capabilities your agent needs: +## Authentication -```sh -agentbricks tools add genie-one --name genie_one --auth user -agentbricks tools add genie-agent SPACE_ID --name genie_agent --auth user -agentbricks tools list --kind genie-one -agentbricks tools list --kind genie-agent -agentbricks tools remove genie_one -agentbricks tools remove genie_agent -``` +Agent Bricks CLI uses [Databricks authentication](https://docs.databricks.com/aws/en/dev-tools/cli/authentication). +`agentbricks login --profile ` validates existing credentials and, if needed in an +interactive terminal, runs `databricks auth login`. It remembers the selected profile in +`~/.agentbricks/config.json`; `agentbricks logout` forgets that selection without revoking its +credentials. For non-interactive use, authenticate the profile first. You can skip `login` when +Databricks SDK default authentication is already configured, or pass `--profile/-p` before an +individual command. Use `--output json` for scripting. -`--name` is optional and defaults to `genie_one` or `genie_agent`, respectively. `--auth` defaults -to `user`; choose `--auth app` deliberately for App service-principal execution. Existing manifests -without `auth` preserve App/default identity. Both add commands and `remove` accept `--source PATH` -to select a project instead of the current directory. Discovery needs no project; read that -project's `agent.toml` to inspect configured bindings. For scripted output, put the global -`-o json` option before `tools`, as in -`agentbricks -o json tools add genie-one --source ./my-agent`. Adding a binding is offline: it updates -`agent.toml` without contacting Genie or checking permissions. The corresponding sources are: - -```toml -[[tools]] -id = "genie_one" -auth = "user" -source = { kind = "genie_one" } - -[[tools]] -id = "genie_agent" -auth = "user" -source = { kind = "genie_agent", space_id = "" } -``` +## How it works -Replace `SPACE_ID` or `` with an existing space's 32-character lowercase hexadecimal -ID. `genie-one` connects to the workspace-wide MCP endpoint -`https:///api/2.0/mcp/genie`, without a space suffix. `genie-agent` uses the -native Genie **Chat-mode** conversation API through the Databricks SDK, not the streaming -Agent-mode API or the per-space MCP endpoint. - -Each native binding exposes `{id}_ask`, `{id}_poll`, and `{id}_query_result`, where `{id}` is its -binding name. Ask accepts an optional `conversation_id` for follow-ups. Ask and poll share a -120-second budget per call, including client setup and submission. If the response is still -running, they return `timed_out` with the conversation and message IDs so the caller can poll -again. If submission times out before a message ID is received, ask returns -`INDETERMINATE_SUBMISSION`: the request may still complete, so do not resubmit automatically. -`NOT_SUBMITTED` means client setup timed out before sending the question. Query results include -the first 100 rows, column schema, a truncation indicator, and a deep link to the conversation. - -Both framework modules, `databricks_agentkit.langgraph` and `databricks_agentkit.openai`, export -`genie_tools()`. New managed-server templates (`server = "agentbricks"`) use it automatically for native Genie Agent bindings; -Genie One uses the existing managed MCP helpers. In an existing project with `server = "agentbricks"`, import -`genie_tools` from your framework module and add `*genie_tools()` to the agent's existing tool list. -The CLI does not patch existing Python code. - -Both paths use Databricks authentication and the routed workspace. Genie One requires the -Managed MCP Servers workspace preview; delegated access requires the `genie` OAuth scope. -The effective caller needs access to the data, the SQL warehouse, and the selected Genie space -where applicable. Agent Bricks CLI does not grant permissions or promise a service-principal fallback when -caller credentials lack access. An offline add succeeding does not establish runtime access. - -The opt-in live tests exercise both frameworks against the configured workspace and an existing -Genie space. From `integrations/agentbricks`, with both framework extras installed: +`agentbricks init` copies framework and runtime files into your project; you own and can edit +those files. The `databricks-agentbricks` dependency supplies AgentKit and the runtime library. +`agentbricks deploy` uses your source and `agent.toml` declarations to run the agent as a +Databricks App. -```sh -DATABRICKS_CONFIG_PROFILE=my-workspace RUN_AGENTBRICKS_GENIE_TESTS=1 \ - AGENTBRICKS_GENIE_SPACE_ID=SPACE_ID \ - uv run pytest tests/integration_tests/genie_tools_test.py -``` +![Deployment: from a local project to a Databricks App](docs/deployment.svg) -By default they ask for the row count of `samples.nyctaxi.trips`. Set `AGENTBRICKS_GENIE_QUESTION` for -another dataset and `AGENTBRICKS_GENIE_EXPECTED_VALUE` to assert a known result cell. +## Continue with a specific task -## Initialize the chat app demo - -The chat app is a LangGraph-specific init overlay, not a command that mutates an existing project. -It is included by default for `--framework langgraph`; pass `--disable-chat-app` to scaffold the -API-only backend instead. +| Task | Guide | +| --- | --- | +| Configure memory, sessions, or the AgentKit SDK | [Memory and sessions](docs/memory-and-sessions.md) | +| Add managed MCP, sandbox, UC function, or Genie tools | [Agent tools](docs/agent-tools.md) | +| Control deployment dependencies and resource lifecycle | [Deployment and lifecycle](docs/deploy-and-maintain.md) | +| Upgrade a customized generated project | [Upgrade a generated project](docs/upgrading-generated-projects.md) | +| Use streaming, background requests, or recovery | [Runtime guide](src/databricks_agentkit/runtime/README.md) | +| Bring an existing agent | [Migration guide](docs/migrating-existing-agents.md) | +| Inspect commands and options | [CLI reference](cli.md) | +| Understand the generated chat UI | [LangGraph](src/databricks_agentbricks/templates/ui/agent-langgraph/CHAT_APP.md) or [OpenAI Agents SDK](src/databricks_agentbricks/templates/ui/agent-openai/CHAT_APP.md) | -```sh -agentbricks init --framework langgraph \ - --profile \ - ./my-agent -cd ./my-agent -agentbricks dev -``` +## Commands -The chat app includes synchronous, SSE streaming, background polling, Session Store, Memory Store, -and HITL resume UI. The framework-specific overlay adds `ui/`, `runtime/ui.py`, the UI-enabled -`runtime/main.py`, and UI tests. +See [the CLI command reference](cli.md) for all commands, arguments, and options. Built-in +examples are available at each level, such as `agentbricks deploy --help`. -For the full deployed demo, bind both managed stores, then deploy: +For zsh completion, add this to `~/.zshrc`: ```sh -agentbricks sessions bind agent-bricks-demo-sessions -agentbricks memory bind agent-bricks-demo-memory -agentbricks --profile deploy agent-bricks-agent-demo --source . +eval "$(_AGENTBRICKS_COMPLETE=zsh_source agentbricks)" ``` -(`bind` declares the store name in `agent.toml`; `agentbricks deploy` creates any declared-but-missing -store and grants the app's service principal access to it. The memory store id flows to the runtime -via the `AGENT_MEMORY_STORE` env var that `deploy` injects; `agentbricks dev` runs locally with memory off -and does not inject it. The id is not persisted in `agent.toml`.) - -The chat UI generates a stable application session UUID in browser local storage, sends it as the -invocation's top-level `session_id`, and creates a fresh invocation UUID per turn. The chat app also -sends this session UUID in the `X-Routing-Key` request header, which is used verbatim to pin -the session to one app replica (it must be non-blank and no more than 128 UTF-8 bytes). The header is neither -authentication nor the template's application session state; it is independent sticky-routing -plumbing. - -The generated `README.md` documents every request the client makes: config discovery, sync and SSE -invocations, background submission and polling, session transcript loading, HITL resume, and memory -entry operations. Capability colors are automatic from `/api/demo/config`; only the -sync/streaming/background transport selector is manual. - ## Contributing -Developing Agent Bricks CLI (`agentbricks`), AgentKit, the runtime, and templates - plus the local dev loop and how to -test unreleased changes on `agentbricks dev` and `agentbricks deploy`, is covered in -[CONTRIBUTING.md](CONTRIBUTING.md). +To change the CLI, AgentKit, runtime, or templates in this repository, see +[CONTRIBUTING.md](CONTRIBUTING.md) for local setup, testing, and releases. diff --git a/integrations/agentbricks/cli.md b/integrations/agentbricks/cli.md index e6f46af75..15fb026d2 100644 --- a/integrations/agentbricks/cli.md +++ b/integrations/agentbricks/cli.md @@ -22,7 +22,8 @@ pip install databricks-agentbricks The `databricks-agentbricks` distribution provides the `agentbricks` command and AgentKit SDK. -See [Installation](README.md#installation) for installing from source and for shell completion. +See [Installation](README.md#installation) for installing from source and +[Commands](README.md#commands) for shell completion. ## Authentication @@ -1213,6 +1214,7 @@ Invoke arbitrary HTTP endpoints. #### `agentbricks endpoint invoke` Send one HTTP request to a Databricks App or arbitrary URL. +Invoking a deployed App requires an OAuth-authenticated profile; PAT profiles are rejected. ``` agentbricks endpoint invoke [APP] [options] diff --git a/integrations/agentbricks/docs/agent-tools.md b/integrations/agentbricks/docs/agent-tools.md new file mode 100644 index 000000000..43f631cce --- /dev/null +++ b/integrations/agentbricks/docs/agent-tools.md @@ -0,0 +1,351 @@ +# Agent tools + +For projects with `[agent].server = "agentbricks"` (the default from `agentbricks init`), +`agent.toml` declares Databricks-managed tool bindings. `agentbricks tools add` updates +that manifest; direct TOML edits have the same behavior. Both framework adapters read +the bindings at runtime without generating or patching agent source: + +```sh +agentbricks tools add sandbox --scope table:samples.nyctaxi.trips +agentbricks tools add mcp system.ai.web_search +agentbricks tools add uc-function catalog.schema.lookup_ticket +agentbricks tools add genie-one +agentbricks tools add genie-agent SPACE_ID +agentbricks tools remove mcp system.ai.web_search +agentbricks tools list +``` + +## Choose tool identity + +`agentbricks tools add mcp`, `agentbricks tools add sandbox`, `agentbricks tools add genie-one`, and +`agentbricks tools add genie-agent` write explicit `auth = "user"` by default. +Use `--auth app` for the App service principal instead. This field is on the tool entry, not +inside `source` or `policy`: + +```toml +[[tools]] +id = "web_search" +auth = "user" +source = { kind = "mcp", service = "system.ai.web_search" } +``` + +Direct UC-function bindings remain app/default identity and do not accept `--auth user`. +Managed-tool add commands write the selected identity to `agent.toml`; inspect that manifest to +review configured bindings. Missing legacy auth continues to mean App identity at runtime; it is +never silently upgraded to user identity. + +## Automatic App-identity access on deploy + +`agentbricks deploy` reconciles least-privilege access for resources explicitly declared by App/default +identity tool bindings. It skips every `auth = "user"` binding because those calls use the request +user's permissions instead of the App service principal. + +| Explicit `agent.toml` resource | Automatic App service-principal access | +| --- | --- | +| UC function | Apps `uc_securable`: `FUNCTION` / `EXECUTE` | +| Genie Agent space | Apps `genie_space`: `CAN_RUN` | +| Sandbox table scope | Apps `uc_securable`: `TABLE` / `SELECT` or `MODIFY` | +| Sandbox volume scope | Apps `uc_securable`: `VOLUME` / `READ_VOLUME` or `WRITE_VOLUME` | +| Sandbox Workspace path | Workspace ACL: `CAN_READ` or `CAN_EDIT` | +| External MCP service | Unity Catalog: effective `EXECUTE` plus `USE_SCHEMA` and `USE_CATALOG` on its named parents | +| Built-in `system.ai` MCP service, including Sandbox and Genie One | Platform-managed access defaults; Agent Bricks does not mutate system securables | + +Native Genie One has no resource identifier in its binding, so it does not add a resource-specific +grant. Use a Genie Agent binding when the App identity should be scoped to one explicit Genie Space. + +Apps-backed tool resources are named deterministically and reconciled to the manifest on each +deploy: removing a binding removes that Agent Bricks-owned Apps resource while preserving Runtime +Store, tracing, and user-owned resources. MCP and Workspace ACL grants are additive in this release +because their permission APIs do not expose trustworthy Agent Bricks ownership metadata; removing +those bindings does not revoke an independently valid grant. + +Only direct resources are automatic. Agent Bricks never discovers or grants tables and warehouses +used by a Genie Space, objects called by a UC function, or resources wrapped by an MCP service. +Grant those transitive dependencies manually when the called service uses the App identity. UC and +Workspace grant checks, plus initial App-resource attachment, happen before source upload. Final +Agent Bricks-owned App-resource reconciliation runs after rollout and can fail after the source has +been uploaded. + +## Request-user authentication + +`DurableAgentServer` derives its request-auth policy directly from the managed tool bindings in +`agent.toml`. Projects do not maintain a separate request-auth contract marker: a managed tool with +`auth = "user"` requires a transient request-user credential. Code-first tools can declare any +additional API scopes that Agent Bricks cannot infer from Python: + +```toml +[auth.user] +required = true +additional_api_scopes = ["sql"] +``` + +`additional_api_scopes` is additive: deploy unions it with scopes inferred from managed bindings, +deduplicates the result, and preserves unrelated scopes already configured on the App. Scope names +are not restricted to a client-side allowlist; Databricks Apps validates whether a requested scope +is supported. Entries must be non-empty strings without surrounding whitespace or control +characters, and a non-empty list requires `required = true`. This request-auth contract is supported +only with `[agent].server = "agentbricks"`. + +Generated framework adapters pass the request-bound resolver to agent construction. A code-first +tool should obtain its user client from that resolver inside the active invocation rather than +creating or persisting a user credential: + +```python +from langchain_core.tools import tool + + +def sql_tools(workspace_client_for): + @tool + def run_statement(statement: str) -> str: + client = workspace_client_for("user") + response = client.statement_execution.execute_statement( + warehouse_id="...", + statement=statement, + ) + return str(response.result) + + return [run_statement] +``` + +The resolver is request-bound and closes after the attempt. Agent Bricks does not inject it into +arbitrary auto-discovered decorated tools; build those tools from the resolver passed to the +generated request-aware agent function. + +Request-user invocations use the same synchronous, streaming, background, status, event-replay, +and idempotency APIs as app-auth invocations. The Runtime Store records only token-free request +state, events, and results. The forwarded credential stays process-local for the active first +attempt and closes when that attempt completes, fails, or is cancelled. A replacement attempt after +failure recovery stops with `MCP_USER_AUTH_RECOVERY_UNSUPPORTED` because no user credential is +available; neither the invoke nor recovery handler runs for that attempt. + +Before deploying user-auth tools from an older project, migrate its request handler and framework +adapter to the current request-auth-aware `DurableAgentServer` template, then explicitly choose `user` or +`app` on **every** managed MCP, sandbox, or Genie entry. Changing `agent.toml` alone does not +upgrade copied Python adapter code. Outdated adapters fail closed rather than silently using App +identity. App-only legacy projects and generic bring-your-own source directories keep the existing +path. + +### Apps user scopes + +Deploy derives Apps user scopes from explicit `auth = "user"` bindings and unions them with +`[auth.user].additional_api_scopes`: + +| Binding | Requested Apps scopes | +| --- | --- | +| Managed MCP (governed ingress) | `ai-gateway` | +| `system.ai.dbsql` | `ai-gateway`, `sql` | +| `system.ai.genie_one_mcp` | `ai-gateway`, `genie` | +| Sandbox with a Volume downscope | `ai-gateway`, `files` | +| First-class Genie One or Genie Agent | `genie` | +| Sandbox with token injection | `ai-gateway`, `workspace.workspace` | +| Sandbox with token injection disabled | `ai-gateway` | + +For example, bind Genie tools in a current project with `server = "agentbricks"`: + +```sh +agentbricks tools add mcp system.ai.genie_one_mcp --auth user +agentbricks tools add genie-agent SPACE_ID --auth user +``` + +Mixed bindings and explicit additions request the union. App-auth and legacy bindings add no user +scopes. These are +explicit service-consent scopes, not a claim that gateway access alone authorizes the downstream +resource. OAuth consent does not grant Unity Catalog privileges: the user still needs access to +the configured Genie Space and its underlying data. + +For a new App, deploy explicitly enables user-token forwarding and includes these scopes in the +initial typed SDK create request before uploading source. An existing App that is missing a required +scope needs one-time explicit permission: + +```sh +agentbricks --profile my-workspace deploy my-agent --allow-user-scope-update +``` + +The `system.ai.dbsql` managed MCP additionally requests the Apps `sql` user scope. This is full SQL +API consent, not `sql:restricted-query`; read-only enforcement remains the service policy plus the +requesting user's Unity Catalog grants. DBSQL does not use Databricks Connect. + +When a user-auth sandbox has `databricks_access_token_included = true`, it requests the Apps +`workspace.workspace` user scope so the injected credential can call workspace APIs. A sandbox +binding with a Volume downscope additionally requests the Apps `files` user scope. +OAuth consent does not grant Volume access: the requesting user still needs the corresponding +Unity Catalog privileges, and the sandbox downscope remains authoritative. A sandbox binding with +token injection disabled does not request `workspace.workspace`; its other resource-derived scopes +still apply. Databricks Apps rejects the legacy bare `workspace` scope, so Agent Bricks requests +`workspace.workspace`. These scopes are requested only for `auth = "user"`; `auth = "app"` uses +the App service principal's permissions instead. + +Review the target App's scopes and coordinate with its other owners before allowing the update. Once +those scopes are present, later deploys do not need the flag. The CLI preserves unrelated scopes, +updates only user scopes and any explicitly requested instance counts, and checks requested **and +effective** scopes before source rollout. It checks for scope changes since preflight, but Apps +read/write is **not atomic**; this is not a lock or a compare-and-swap guarantee. Polling is bounded +and a mismatch stops source deployment. +Users may need to sign out and **re-consent** after changing scopes; effective-scope verification +does not refresh an existing user's consent. + +Apps may report `iam.access-control:read` and `iam.current-user:read` as implicit effective +scopes. The CLI permits these platform defaults during verification but does not request them +as configurable scopes. Explicitly disabled user-token forwarding stops deployment; enable +forwarding and restart the App compute before retrying. + +Removing a tool or switching back to app-only auth **does not remove Apps scopes**. Remove +unneeded scopes explicitly in Databricks Apps, and verify both configured and effective scopes +before declaring removal complete. The CLI does not send empty-list scope updates: the SDK's +`App.as_dict()` omits empty lists, so that would not prove removal succeeded. No scopes are +managed for generic bring-your-own apps without this managed user contract. + +App-auth tools execute with workload privileges. Restrict App `CAN USE` to callers trusted +for **all** App-auth tools, or deploy those tools separately. Models, custom MCP servers, +Memory/Session Stores, and tracing keep their existing credentials. + +## Discover and manage bindings + +For MCP services, the remove command accepts the same service name as the add command. You can also +remove any binding by its `id` in `agent.toml`, for example `agentbricks tools remove web_search`. +Every successful add (including an already-configured no-op) points you to the target project's +`agent.toml` to review configured managed tools and MCP bindings. With `--source`, the message +points to that project's file. JSON add output includes its path in `manifest`. + +`agentbricks tools list` discovers **available integrations to add**, not configured bindings. By default +it shows built-in add recipes and caller-visible MCP Services in `system.ai`. A recipe may still +need your resources: sandbox scopes, a concrete UC function name, or a Genie Space ID. Genie One +needs no additional argument. `system.ai.sandbox` is represented by its scoped recipe rather than +a second unscoped add command. The list does not enumerate every workspace schema, individual +operations inside MCP services, or custom Python tools. + +`agentbricks tools add mcp` looks up the service in the selected workspace before writing `agent.toml`. +Use `agentbricks --profile tools add mcp ` to select a workspace. A missing service or +failed lookup (including authentication or permission errors) leaves the project unchanged. This +checks service metadata access, not whether every tool can be executed at runtime. Removing local +bindings does not require workspace access. + +```sh +agentbricks tools list +agentbricks tools list --kind mcp +agentbricks tools list --kind mcp --schema main.tools +agentbricks tools list --kind sandbox +agentbricks tools list --kind genie-one +agentbricks tools list --kind genie-agent +agentbricks --output json tools list +``` + +No agent project is required for discovery. MCP discovery uses your Databricks profile; the +`sandbox`, `uc-function`, `genie-one`, and `genie-agent` kind filters show local recipes without +authentication. `--schema` requires `--kind mcp` and replaces the default `system.ai` scope. An +API/authentication failure returns nonzero and marks discovery incomplete, while retaining local +recipes; it is not reported as an empty successful discovery. Listing metadata does not verify +runtime execution permissions. + +**Migration:** the former configured `tools list` view and its `--source` option are removed. +Read `agent.toml` (its `[[tools]]` entries) to inspect configured bindings. Discovery JSON uses +`schema_version: 2`, with `available_tools` (`name`, `kind`, `add_command`), `mcp_schema` (null for +local-only recipes), `complete`, and `errors`. Replace old scripts that read configured-list JSON +with TOML inspection. Replace `agentbricks mcp list [--schema catalog.schema]` with +`agentbricks tools list --kind mcp [--schema catalog.schema]`; the former command is removed. Use +`agentbricks tools list --help` for the new discovery contract. + +## Custom Python tools + +In managed-server templates, custom Python tools are code-first. Write them with the framework's native +decorator in `agent/tools/`: LangGraph uses `@tool`, while OpenAI Agents uses `@function_tool`. The +templates auto-discover decorated tools from that package and add them to the agent; there is no CLI +command or `agent.toml` entry to keep in sync. Customer-managed MCP servers are likewise ordinary +code in `agent/mcps.py` and are joined with the managed bindings by `mcp_tools(...)` or +`mcp_servers(...)`. See the [first customization example](../README.md#first-customization-add-a-custom-tool) +for a working example in both frameworks. + +Projects created with `--server custom` do not auto-discover `agent/tools/` or load managed tool +bindings from `agent.toml`, so `agentbricks tools add` rejects those projects. Wire framework-native Python +tools and MCP servers directly in `agent/agent.py` instead. + +If a manifest with `server = "agentbricks"` contains `source = { kind = "python", ... }`, remove that +`[[tools]]` entry; the decorated tool in `agent/tools/` remains active. `agentbricks dev` and `agentbricks deploy` +do not generate or patch Python tool code, and do not alter the manifest's `[[tools]]` bindings. + +## Sandbox policy + +Sandbox scopes default to read-only access. Repeat `--scope` to allow more than one resource, use +`volume:` or `workspace:` for those resource types, and use `--permission read_write` only when the +agent needs writes. Every sandbox call carries this fixed downscope in MCP `_meta`, outside the tool +arguments controlled by the model. New sandbox bindings also expose the selected Databricks +credential to sandbox code by default: + +```toml +[[tools]] +id = "sandbox" +auth = "user" +source = { kind = "sandbox", service = "system.ai.sandbox" } +policy = { downscope = [{ resource = "workspace:/Workspace/Shared", permission = "read_only" }], databricks_access_token_included = true } +``` + +With `databricks_access_token_included = true`, the sandbox receives `DATABRICKS_HOST`, a short-lived +`DATABRICKS_TOKEN`, and `DATABRICKS_AUTH_TYPE`, so code such as +`WorkspaceClient().current_user.me()` can call workspace APIs. This policy does not choose the +identity: `auth = "user"` uses the request user's OBO credential, while `auth = "app"` uses the +Databricks App service principal. Use `--no-databricks-access-token-included` when adding a sandbox that +does not need workspace API access. Existing manifests that omit `databricks_access_token_included` +remain disabled until explicitly updated. + +## Genie tools + +Genie One and Genie Agent support ship with Agent Bricks CLI, but bindings are opt-in, like sandbox tools. +Installing the current `databricks-agentbricks` distribution does not configure a Genie Space ID or enable a Genie binding. Add only the +capabilities your agent needs: + +```sh +agentbricks tools add genie-one --name genie_one --auth user +agentbricks tools add genie-agent SPACE_ID --name genie_agent --auth user +agentbricks tools list --kind genie-one +agentbricks tools list --kind genie-agent +agentbricks tools remove genie_one +agentbricks tools remove genie_agent +``` + +`--name` is optional and defaults to `genie_one` or `genie_agent`, respectively. `--auth` defaults +to `user`; choose `--auth app` deliberately for App service-principal execution. Existing manifests +without `auth` preserve App/default identity. Both add commands and `remove` accept `--source PATH` +to select a project instead of the current directory. Discovery needs no project; read that +project's `agent.toml` to inspect configured bindings. For scripted output, put the global +`-o json` option before `tools`, as in +`agentbricks -o json tools add genie-one --source ./my-agent`. Adding a binding is offline: it updates +`agent.toml` without contacting Genie or checking permissions. The corresponding sources are: + +```toml +[[tools]] +id = "genie_one" +auth = "user" +source = { kind = "genie_one" } + +[[tools]] +id = "genie_agent" +auth = "user" +source = { kind = "genie_agent", space_id = "" } +``` + +Replace `SPACE_ID` or `` with an existing space's 32-character lowercase hexadecimal +ID. `genie-one` connects to the workspace-wide MCP endpoint +`https:///api/2.0/mcp/genie`, without a space suffix. `genie-agent` uses the +native Genie **Chat-mode** conversation API through the Databricks SDK, not the streaming +Agent-mode API or the per-space MCP endpoint. + +Each native binding exposes `{id}_ask`, `{id}_poll`, and `{id}_query_result`, where `{id}` is its +binding name. Ask accepts an optional `conversation_id` for follow-ups. Ask and poll share a +120-second budget per call, including client setup and submission. If the response is still +running, they return `timed_out` with the conversation and message IDs so the caller can poll +again. If submission times out before a message ID is received, ask returns +`INDETERMINATE_SUBMISSION`: the request may still complete, so do not resubmit automatically. +`NOT_SUBMITTED` means client setup timed out before sending the question. Query results include +the first 100 rows, column schema, a truncation indicator, and a deep link to the conversation. + +Both framework modules, `databricks_agentkit.langgraph` and `databricks_agentkit.openai`, export +`genie_tools()`. New managed-server templates (`server = "agentbricks"`) use it automatically for native Genie Agent bindings; +Genie One uses the existing managed MCP helpers. In an existing project with `server = "agentbricks"`, import +`genie_tools` from your framework module and add `*genie_tools()` to the agent's existing tool list. +The CLI does not patch existing Python code. + +Both paths use Databricks authentication and the routed workspace. Genie One requires the +Managed MCP Servers workspace preview; delegated access requires the `genie` OAuth scope. +The effective caller needs access to the data, the SQL warehouse, and the selected Genie space +where applicable. Agent Bricks CLI does not grant permissions or promise a service-principal fallback when +caller credentials lack access. An offline add succeeding does not establish runtime access. diff --git a/integrations/agentbricks/docs/deploy-and-maintain.md b/integrations/agentbricks/docs/deploy-and-maintain.md new file mode 100644 index 000000000..a38649224 --- /dev/null +++ b/integrations/agentbricks/docs/deploy-and-maintain.md @@ -0,0 +1,82 @@ +# Deploy and manage a generated agent + +This guide is for owners of projects generated by `agentbricks init`. It explains which files and +dependency inputs control a deployment, how application resources and agent state behave across local +runs, redeployments, and deletion. For command syntax and options, see the [CLI command reference](../cli.md). +To upgrade a customized scaffold, see [Upgrade a generated agent project](upgrading-generated-projects.md). + +## Deployment dependency inputs + +The generated `pyproject.toml` declares Python packages; `app.yaml` runs `uv run start-server` in +Databricks Apps. Released packages can come from the configured package index (public PyPI by +default, or `agentbricks deploy --pip-index-url `). For the advanced case of deploying an +unreleased package, pin a pushed Git commit that the Apps build can reach: + +```toml +[tool.uv.sources] +databricks-agentbricks = { git = "https://github.com//databricks-ai-bridge", rev = "", subdirectory = "integrations/agentbricks" } +``` + +Commit and push the ref before deploying; the Apps build clones that commit. A local path or +`file://` source is unavailable inside the Apps build. Contributor-only local-path testing is +documented in [CONTRIBUTING.md](../CONTRIBUTING.md#testing-sdk--runtime-changes-in-a-scaffold). + +`agentbricks deploy` uploads the project source but excludes the local `uv.lock`; the Apps build +resolves dependencies against its own index. The generated `>=` requirements can therefore resolve +to newer packages on a later deployment. Pin direct dependency versions in `pyproject.toml`, verify +the selected package index and reachable Git commits, then compare the local environment with the +deployed build logs. The current deploy flow does not provide a frozen transitive dependency graph +from the local lockfile. Keep `agent.toml` bindings and the chosen app name alongside the dependency +manifest so the same deployment targets the same managed resources. + +## Resource and state lifecycle + +The default managed-server template declares memory, session, and tracing names in `agent.toml`. +`init` writes those declarations without creating workspace resources. `deploy` resolves them in the +target workspace, creates missing resources, reuses accessible ones with matching names, creates or +updates the App, rolls out the source, then attempts the App's store and trace access grants. An existing +name that the caller cannot access causes an error. A grant failure can leave a deployed App without the +corresponding feature; inspect deploy warnings. +`agentbricks memory/sessions bind` and `unbind` edit `agent.toml`; they do not delete remote stores. +After a memory or session unbind, redeploy currently leaves any earlier `AGENT_MEMORY_STORE` or +`AGENT_SESSION_STORE` setting in `app.yaml` in place. Remove the stale setting from `app.yaml` before +redeploying if you want the App to stop using that store. A clean tracing unbind is removed on the +next deploy. +Default store names contain a six-letter token (`--memory` and +`--sessions`); use `agentbricks memory bind ` or +`agentbricks sessions bind ` to select existing stores. The default tracing experiment is +under `/Shared/agentbricks_traces/`; `agentbricks tracing list` shows available traces. + +The Invocation Runtime Store row in the matrix applies only to projects with +`[agent].server = "agentbricks"`. A custom server defines its own protocol and does not receive a +managed Runtime Store. + +| Resource or state | Created or reused | Local `dev` and restart | Redeploy and cleanup | +| --- | --- | --- | --- | +| Project files and dependencies | `init` copies a template; `dev` builds `.venv` from `pyproject.toml`. | Source files stay on disk. The local environment is reused until `dev --prepare-environment` rebuilds it. | Deploy syncs source and resolves dependencies again without the local `uv.lock`. App deletion leaves the local project alone. | +| Databricks App | Deploy creates the named App or reuses it, then updates compute and source. | `dev` serves the project locally without creating an App. | Redeploy with the same name updates that App; `agentbricks deployments delete ` deletes it. | +| Invocation Runtime Store | The current default deploy creates or reuses an App-owned database in the workspace's shared Lakebase project. An internal legacy path uses a per-App Lakebase project. | `dev` keeps invocation status, results, and events in process; they disappear on restart. | Deployed invocation records persist across restart and redeploy. Queued work can resume; active work needs a recovery handler and may run more than once. `agentbricks deployments delete` removes the managed store before the App; the legacy delete path does not explicitly remove its Lakebase project. | +| Managed tool access | Deploy reconciles direct App-auth tool grants from `agent.toml` before source upload, then finalizes Agent Bricks-owned App resources after rollout; request-user tools use the caller's permissions. | `dev` creates no App service principal or Apps grants. | Removing a tool binding removes Agent Bricks-owned Apps resources on redeploy. MCP and Workspace grants are additive; see [automatic App-identity access](agent-tools.md#automatic-app-identity-access-on-deploy). | +| Memory Store | Deploy creates a declared store if missing or reuses an accessible store by name, then attempts the App grant. | Managed long-term memory is off in `dev`. | Memory persists independently of the App; redeploy reuses the bound store. Unbinding or deleting the App does not delete it. Remove the stale `app.yaml` setting to detach it after unbind; use the separate store delete command when appropriate. | +| Session Store | Deploy creates or reuses a declared store by name, then attempts the App grant. | `dev` keeps conversation state in process, so a restart loses it. | A bound store preserves LangGraph checkpoints and OpenAI Agents SDK transcripts across restart and redeploy. OpenAI pending approval `RunState` stays in process. Unbinding or deleting the App does not delete the store; remove the stale `app.yaml` setting to detach it. | +| MLflow traces | `dev` uses a local MLflow server; deploy attempts to create or reuse the bound workspace experiment and grant App access. | Local traces are recorded in `.agentbricks/`; they remain on disk after `dev` stops. | Deployed traces remain in the workspace experiment. Unbinding removes the App's tracing configuration on a later clean deploy; it does not delete the experiment. | + +Redeploying an older App can attach the current managed Runtime Store without migrating invocation +records from its legacy per-App Lakebase project. Managed-store cleanup errors retain the App for +retry; deleting the App directly bypasses that cleanup. + +The Runtime Store tracks HTTP invocations, status, results, and event replay. The framework's +conversation history belongs to its Session Store when bound. The generated chat UI keeps its +session ID in browser local storage and sends that ID as the top-level `session_id` with each turn. +`DurableAgentServer` accepts an optional top-level `session_id` for generic handlers, but the generated +LangGraph and OpenAI Agents templates require a nonempty top-level `session_id` on every invocation. +API clients should reuse that value for conversation continuity and send a new invocation `id` for each +turn. +By default, the templates use the session ID as the state actor; request-user-authenticated +invocations namespace session state by user. LangGraph can resume from a matching checkpoint after +worker loss; the OpenAI Agents SDK template replays the input against its saved transcript. +Recovery can repeat external side effects, so make tools idempotent. For request-user-authenticated +tools, credentials are not persisted and background recovery is unsupported. + +The default chat UI keeps its session ID across page reloads; see [invoke HTTP endpoints](../README.md#invoke-from-the-command-line) +for the explicit API request shape. diff --git a/integrations/agentbricks/docs/memory-and-sessions.md b/integrations/agentbricks/docs/memory-and-sessions.md new file mode 100644 index 000000000..92f3dc4dd --- /dev/null +++ b/integrations/agentbricks/docs/memory-and-sessions.md @@ -0,0 +1,203 @@ +# Memory and sessions + +An agent needs two kinds of state: the state of the interaction it is handling right now, and the +durable knowledge it carries from one conversation to the next. Databricks provides a managed store +for each, both backed by Lakebase and usable from agents built on any framework: + +- **Managed agent sessions** store an agent's session state for one interaction. Most commonly this + is the conversation history — the ordered transcript of messages, tool calls, and results — but it + can be any state a framework persists, such as a LangGraph graph. The agent reads it at the start + of a turn and appends to it as the interaction runs. +- **Managed agent memory** stores durable facts, preferences, and decisions that an agent recalls in + later, separate conversations through text search. + +![Sessions and memory: the agent reads and appends one conversation's transcript in the session store, and recalls and saves durable facts in the memory store, which outlive any single conversation.](sessions_and_memory.png) + +The examples below use the [`AgentKitClient` Python SDK](#agentkit-sdk). The same operations are +available as `agentbricks sessions` and `agentbricks memory` CLI commands; see the [CLI command +reference](../cli.md) for the complete command set. + +## AgentKit SDK + +Use `AgentKitClient` from `databricks_agentkit` with an authenticated Databricks +`WorkspaceClient`, or omit the client to use the default Databricks SDK authentication. Its +`memory_stores` and `session_stores` collections create, get, and list stores. A returned store +manages its entries or sessions; returned memories and sessions own their `update()` and +`delete()` operations. + +All `list()` methods return iterators that consume server pages automatically. List `page_size` and +search `limit` values must be between 1 and 100. `session.list_items()` also auto-pages. + +## Sessions + +A **session store** holds **sessions**, and each session holds an ordered list of **session items**. +A session is one interaction — typically a conversation thread — grouped under an `actor_id` (who it +belongs to; set this from trusted application context, never a model- or user-supplied value) and +identified by a caller-chosen `session_id` (the service generates one if you omit it). Each item is +an opaque, JSON-compatible `data` value — a message, tool call, result, or reasoning block — that +Databricks stores and returns verbatim, in order, and never mutates once appended. + +Create a store, start a session, append the conversation's turns, and read the history back on a later +request: + +```python +from databricks.sdk import WorkspaceClient +from databricks_agentkit import AgentKitClient + +agentkit = AgentKitClient(WorkspaceClient()) + +session_store = agentkit.session_stores.create("support-agent-sessions") +session = session_store.add(actor_id="customer-123", session_id="case-456") + +session.append_items( + [ + {"type": "message", "role": "user", "content": "I need help with my cluster."}, + {"type": "message", "role": "assistant", "content": "Let's take a look."}, + ] +) + +# On a later turn, reload the session and read its full history in order. +session = session_store.get("case-456") +history = [item.data for item in session.list_items()] # list_items auto-pages +``` + +A session can be **forked** into an independent branch: a new session seeded with the original's +history, linked back to its origin by `parent_session_id`. Fork the full history, or only up to a +specific item, to explore an alternate continuation without disturbing the original thread: + +```python +branch = session.fork(actor_id="customer-123") # add up_to_item_id=... to branch up to one item +``` + +Deleting a session that has such descendants requires `session.delete(force=True)` to cascade. + +## Memory + +A **memory store** holds **memory entries**. Each entry is a free-form `content` string plus a short +`description` used for retrieval, keyed by three fields: `actor_id` (whose memory it is — set from +trusted application context, never a model- or user-supplied value), `path` (a filesystem-like key +within an actor, such as `/preferences/response-style.md`), and an optional `session_id` (the session +an entry came from, for provenance). An entry is uniquely identified by its `actor_id`, `path`, and +optional `session_id`. + +Write an entry when the agent learns something durable, then recall it in a later, separate +conversation with a natural-language search. Results are ranked by full-text (BM25) relevance, up to +100 entries, with no pagination or vector similarity: + +```python +from databricks.sdk import WorkspaceClient +from databricks_agentkit import AgentKitClient + +agentkit = AgentKitClient(WorkspaceClient()) + +memory_store = agentkit.memory_stores.create("support-agent-memory") +memory_store.add( + actor_id="user-123", + path="/preferences/communication.md", + content="Prefers email over phone. Timezone: PST.", + description="User 123 communication preferences", +) + +# In a later, separate conversation, recall what the agent knows about this user. +results = memory_store.search(actor_id="user-123", query="communication preferences", limit=10) +``` + +To browse rather than search, `memory_store.list(actor_id=..., path_prefix=...)` returns entries +directly. + +> **`actor_id` partitions data; it is not access control.** Both stores are workspace-scoped and +> authorized at the store level, so any principal that can reach a store can read and write every +> actor's entries. For strict isolation between tenants or users, use a separate store per boundary. +> Grant another principal — such as your app's service principal — access with +> `session_store.grant_permission(principal_id)` or `memory_store.grant_permission(principal_id)`; +> `agentbricks deploy` attempts this grant for the deployed app. + +## Framework adapters + +In an agent configured with `server = "agentbricks"`, the framework adapter reads and appends +session state for you. The adapter resolves a bound store when one is configured. During local +`agentbricks dev`, sessions use process-local state and long-term memory is off; a deployed agent +gets the managed stores through the environment configured by `agentbricks deploy`. + +### LangGraph + +Pass `checkpointer()` when you build the agent and scope each run with `thread_config(session_id)`. +Pass the trusted actor as the second argument when the graph's state must be partitioned by actor. + +```python +from databricks_agentkit.langgraph import checkpointer, thread_config + +agent = create_agent(model=..., tools=[...], checkpointer=checkpointer()) +result = await agent.ainvoke(inputs, config=thread_config(session_id)) +``` + +Add the memory tools to the model and execution tool list. `memory_tools(actor)` exposes `remember` +and `recall` bound to one actor's partition. It resolves the store from the `[memory_store]` binding, +carried to the runtime by the `AGENT_MEMORY_STORE` environment variable that `agentbricks deploy` +injects. It returns no tools when no store is set, so the agent runs unchanged: + +```python +from databricks_agentkit.langgraph import memory_tools + +agent = create_agent(model=..., tools=[*your_tools, *memory_tools(actor)]) +``` + +### OpenAI Agents SDK + +The OpenAI Agents adapter exposes the same session and memory capabilities. Pass +`session_store(session_id)` to `Runner.run` and add `memory_tools(actor)` to the agent's tools. The +optional `actor` argument partitions durable session and memory data; derive it from trusted +application context. + +```python +from agents import Agent, Runner +from databricks_agentkit.openai import memory_tools, session_store + +agent = Agent(model=..., tools=[*your_tools, *memory_tools(actor)]) +result = await Runner.run( + agent, + messages, + session=session_store(session_id, actor), +) +``` + +Both adapters resolve their store from the explicit argument first, then the corresponding +`AGENT_SESSION_STORE` or `AGENT_MEMORY_STORE` environment variable, and then the `agent.toml` +binding. With no managed store configured, the session helpers keep state in process and the memory +helper returns no tools. This is the local development behavior; deployment supplies the managed +store configuration. + +## Declaring and provisioning stores + +For a deployed agent, `agent.toml` declares which stores it uses and `agentbricks deploy` provisions +them — you do not create stores by hand for the deployed project. `agentbricks init` declares a +default memory and session store named from the project; override those names, point at stores you +already have, or let `deploy` create them: + +To scaffold a project with default memory and session stores declared in `agent.toml`: + +```sh +agentbricks init my-agent +``` + +To scaffold with specific store names instead: + +```sh +agentbricks init my-agent --memory-store support-agent-memory --session-store support-agent-sessions +``` + +For an existing project, bind the stores in `agent.toml`, then deploy. Binding edits the manifest; +it does not create the remote stores: + +```sh +# Run these commands from the project directory. +cd my-agent +agentbricks sessions bind support-agent-sessions +agentbricks memory bind support-agent-memory + +# deploy creates any declared-but-missing store and attempts the App service-principal grant. +agentbricks deploy my-agent +``` + +Memory and session stores are independent resources: deleting one never affects the other. For runtime +hooks, endpoint behavior, and recovery, see the [runtime guide](../src/databricks_agentkit/runtime/README.md). diff --git a/integrations/agentbricks/docs/migrating-existing-agents.md b/integrations/agentbricks/docs/migrating-existing-agents.md new file mode 100644 index 000000000..20796f8fc --- /dev/null +++ b/integrations/agentbricks/docs/migrating-existing-agents.md @@ -0,0 +1,62 @@ +# Bring an existing agent + +From the existing project, choose the command for your agent's framework to prepare +the migration for your coding agent. + +LangGraph: + +```sh +agentbricks init --framework langgraph --existing . +``` + +OpenAI Agents SDK: + +```sh +agentbricks init --framework openai --existing . +``` + +Before or after the conversion, inspect its progress without changing the repository or contacting +Databricks: + +```sh +agentbricks doctor . +agentbricks -o json doctor . +``` + +Doctor exits 0 only when the project has a valid Agent Bricks manifest and matching project +metadata, uses the Agent Bricks server, declares the framework-appropriate `databricks-agentbricks` +extra and a non-empty `app.yaml` command, constructs `DurableAgentServer` with an `invoke` hook, and +calls a recognized adapter for the selected framework in production Python source. Test, example, and +old/stale directories do not count as source evidence. A failed report is the normal result for a +project that still needs migration; run +`agentbricks init --framework --existing ` with the appropriate framework +to prepare the migration instructions. Doctor never imports or executes the target's source, and a +bounded source scan that exceeds a limit is reported while the evidence it already found still counts. +Its findings are static repository evidence, not proof that the configured startup command executes +the files it finds. + +This writes `agent-bricks-migrate/` containing a skill, a prompt to paste into your coding agent, +`references/migration.json`, and a reference project generated from the templates bundled with the +installed CLI. The bundle sits outside any single agent's configuration directory; `.claude/skills/` +and `.agent/skills/` each receive a small skill that points at it, so Claude Code, Codex, and +similar tools discover the same instructions without duplicating the reference. Agent Bricks CLI +prepares the instructions; the coding agent performs and verifies the conversion. Init leaves application +source, dependencies, `.env`, and existing `.agentbricks/project.toml` configuration intact and refuses to overwrite +existing migration files. + +The bundle is scaffolding for the migration, not part of the application: delete `agent-bricks-migrate/` +and the two pointer skills once the conversion is done, and keep them out of commits meanwhile. + +The skill follows the shared managed-runtime contract for the selected framework, included in new projects and +migration references: +[LangGraph](../src/databricks_agentbricks/templates/agent-langgraph/AGENTKIT_CONTRACT.md) or +[OpenAI Agents SDK](../src/databricks_agentbricks/templates/agent-openai/AGENTKIT_CONTRACT.md). It explicitly +handles existing history, custom state and output, recovery, and client/session contracts. For +LangGraph, switching checkpointers does not migrate old conversations (likewise, the OpenAI Agents +SDK keeps prior Session transcripts and RunState behind); unresolved transitions require a user +decision. + +The reference honors `--disable-chat-app`, `--memory-store`, `--session-store`, and the selected +profile. These are migration intent; init does not provision resources or change the existing +application. Migration supports LangGraph and the OpenAI Agents SDK with the managed server (`server = "agentbricks"`); +`--server custom` is not supported for `--existing`. diff --git a/integrations/agentbricks/docs/upgrading-generated-projects.md b/integrations/agentbricks/docs/upgrading-generated-projects.md new file mode 100644 index 000000000..6503bc063 --- /dev/null +++ b/integrations/agentbricks/docs/upgrading-generated-projects.md @@ -0,0 +1,41 @@ +# Upgrade a generated agent project + +This guide is for owners of app projects generated by `agentbricks init`. It explains which copied +files you own, how to identify the installed runtime, and how to adopt CLI and template releases while +preserving your customizations. For deployment inputs and resource/state behavior, see +[Deploy and manage a generated agent](deploy-and-maintain.md). + +## Project ownership and upgrades + +`agentbricks init` copies the template bundled with the installed CLI into your project. You own the +copied `agent/`, `runtime/`, `app.yaml`, and `pyproject.toml` files; edit them to customize the agent. +The installed `databricks-agentbricks` dependency supplies `databricks_agentkit`, including +`DurableAgentServer` and the framework adapters imported by those files. Updating that dependency +updates the library code. New template files are copied only into new projects, so review and merge +later template changes into a customized project yourself. + +`agentbricks --version` shows the installed CLI version, and `init` prints the bundled template's +package version as `Template ref`. Save that output if you need the template's exact origin: +`.agentbricks/project.toml` records the framework and template name, but not the package version. +The project's `pyproject.toml` declares a version range for the runtime dependency. After the first +`dev` run, check the version actually installed in the project environment with: + +```sh +.venv/bin/python -c "from importlib.metadata import version; print(version('databricks-agentbricks'))" +``` + +To adopt a new release while preserving your changes: + +1. Commit or back up the customized project. Choose the target version and update its + `databricks-agentbricks[langgraph]` or `databricks-agentbricks[openai]` requirement in + `pyproject.toml`. Pin an exact version when you need the same direct dependency on every build. +2. Run `uv lock`, `agentbricks dev --prepare-environment`, and `uv run pytest` from the project. + The explicit environment rebuild is needed because later `dev` runs reuse `.venv`. +3. Upgrade the CLI, scaffold a **different directory** with the same `--framework`, `--server`, + and `--disable-chat-app` choices as your project, and compare its `agent/`, `runtime/`, + `app.yaml`, and `pyproject.toml` with your project. Merge the template changes you want and run + the project tests again. `init` refuses to overwrite an existing directory; it does not upgrade + copied files in place. +4. Redeploy the existing app name and inspect `agentbricks deployments logs ` for the + resolved packages and startup errors. Check the agent through its URL or an + [endpoint invocation](../README.md#invoke-from-the-command-line). diff --git a/integrations/agentbricks/src/databricks_agentkit/runtime/README.md b/integrations/agentbricks/src/databricks_agentkit/runtime/README.md index 53f5e203d..41c5e65dc 100644 --- a/integrations/agentbricks/src/databricks_agentkit/runtime/README.md +++ b/integrations/agentbricks/src/databricks_agentkit/runtime/README.md @@ -4,7 +4,7 @@ Agent Bricks CLI takes your agent code from a local project to a hosted endpoint LangGraph or OpenAI Agents template, or bring an existing agent. - **Deployment:** Scaffold a project, run it locally, and deploy it to Databricks Apps. The CLI - provisions the stores declared in your project, grants the app access, and configures tracing. + provisions the stores declared in your project, attempts the App access grants, and configures tracing. - **Runtime:** `DurableAgentServer` provides synchronous, streaming, and background execution, with persistent results and automatic crash recovery on deployment. Request-user authentication is attached only to the active first attempt and is never persisted. @@ -43,12 +43,12 @@ Use `--framework openai` for OpenAI Agents. Managed-server templates include a c pass `--disable-chat-app` for an API-only project. 1. **Initialize:** `agentbricks init` generates the agent code and runtime adapter separately. It records - the server choice and default `my-agent-memory` / `my-agent-session` bindings in `agent.toml`. + the server choice and distinct default memory and session store names in `agent.toml`. 2. **Develop:** Edit your model, prompts, and tools in `agent/`. `agentbricks dev` runs the project locally. Synchronous, streaming, and background requests use the same Runtime for both authorization policies. -3. **Deploy:** `agentbricks deploy` creates or reuses the declared Session and Memory Stores, grants the - app's service principal access, configures tracing, and deploys the app. For a managed server, +3. **Deploy:** `agentbricks deploy` creates or reuses the declared Session and Memory Stores, attempts + the App's service-principal grants, configures tracing, and deploys the App. For a managed server, it also creates or reuses the deployment's Runtime Store. The generated configuration starts with: @@ -61,15 +61,18 @@ framework = "langgraph" server = "agentbricks" [memory_store] -name = "my-agent-memory" +name = "my-agent-abcdef-memory" [session_store] -name = "my-agent-session" +name = "my-agent-abcdef-sessions" [tracing] -experiment_name = "/Shared/agentbricks_traces/my-agent" +experiment_name = "/Shared/agentbricks_traces/my-agent-abcdef" ``` +Here `abcdef` stands for the six-letter token generated for that project. The same token +appears in its default store and tracing names. + Override store names at initialization with `--memory-store` and `--session-store`, or later with `agentbricks memory bind ` and `agentbricks sessions bind `. Custom-server templates declare these stores only when explicitly requested. Tracing is bound by experiment **name** (its presence turns @@ -116,7 +119,8 @@ forwarded principal so users cannot collide with each other. Clients may also supply an optional top-level `session_id`. Runtime persists it separately from the opaque `input`, serializes invocations that share it, and exposes it as `context.session_id`. If it is omitted, the invocation remains sessionless; Runtime does not infer it from the invocation ID, -the input payload, a handler response, or `X-Routing-Key`. +the input payload, a handler response, or `X-Routing-Key`. The generated LangGraph and OpenAI Agents +adapters require a nonempty top-level `session_id` on every invocation; reuse it across turns. - **Synchronous:** Wait for the result in the POST response. - **Streaming (`stream: true`):** Receive progress events as Server-Sent Events (SSE). @@ -161,12 +165,14 @@ flowchart LR `agentbricks dev` supports the same invocation APIs with an **In-process Runtime Store**. Run state, events, and results are lost when the serving process exits. Interrupted work is not automatically -restarted. Session and Memory Store persistence is separate from this local execution state. +restarted. The generated templates also keep conversation state in process and leave managed +long-term memory off during local development. ### Deployed execution -`agentbricks deploy` provisions a dedicated PostgreSQL database for each deployment with `server = "agentbricks"` and -reuses it on redeployment. Results and events survive worker restarts, and any replica can serve +`agentbricks deploy` provisions an App-owned PostgreSQL database in the workspace's shared Lakebase +project for each deployment with `server = "agentbricks"` and reuses it on redeployment. Results and +events survive worker restarts, and any replica can serve polling and stream-reconnection requests. With a recovery handler registered, the runtime detects stale heartbeats and starts a replacement attempt on an available worker.