Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions integrations/agentbricks/CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,6 +101,32 @@ the same change, in `src/databricks_agentbricks/cli/doctor.py`:
Also refresh the framework references and examples in `cli.md` and `README.md`, and add doctor test
coverage for the new framework in `tests/unit_tests/doctor_test.py`.

## Live tool tests

Read-only tool discovery can be checked against the installed wheel without creating a
project or deploying an agent:

```sh
AGENTBRICKS_E2E_PROFILE=<profile> .venv-functional/bin/pytest tests/e2e/tool_discovery_test.py -v
```

The checks compare default and MCP-filtered discovery with the compatibility service
list. Set `AGENTBRICKS_E2E_SCHEMA=catalog.schema` to exercise another schema. Local
add/review/remove flows and help pages are covered by `tests/functional/cli_smoke_test.py`.

The opt-in Genie tests exercise both frameworks against an existing Genie space. From
`integrations/agentbricks`, with both framework extras installed:

```sh
DATABRICKS_CONFIG_PROFILE=my-workspace RUN_AGENTBRICKS_GENIE_TESTS=1 \
AGENTBRICKS_GENIE_SPACE_ID=SPACE_ID \
uv run pytest tests/integration_tests/genie_tools_test.py
```

By default they ask for the row count of `samples.nyctaxi.trips`. Set
`AGENTBRICKS_GENIE_QUESTION` for another dataset and
`AGENTBRICKS_GENIE_EXPECTED_VALUE` to assert a known result cell.

## Cutting a release

Run **Cut Agent Bricks release** from the Actions tab with a version such as `0.4.0` or
Expand Down
1,002 changes: 85 additions & 917 deletions integrations/agentbricks/README.md

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in general I still and we need to be very hard to follow. I see this as a combination of helping the user to get started and teaching them how to do things, but then there's also a combination of how developers could iterate on this. I feel like the split is not exactly clear. Is there a better way that we can structure?

Large diffs are not rendered by default.

4 changes: 3 additions & 1 deletion integrations/agentbricks/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,8 @@ pip install databricks-agentbricks

The `databricks-agentbricks` distribution provides the `agentbricks` command and AgentKit SDK.

See [Installation](README.md#installation) for installing from source and for shell completion.
See [Installation](README.md#installation) for installing from source and
[Commands](README.md#commands) for shell completion.

## Authentication

Expand Down Expand Up @@ -1213,6 +1214,7 @@ Invoke arbitrary HTTP endpoints.
#### `agentbricks endpoint invoke`

Send one HTTP request to a Databricks App or arbitrary URL.
Invoking a deployed App requires an OAuth-authenticated profile; PAT profiles are rejected.

```
agentbricks endpoint invoke [APP] [options]
Expand Down
351 changes: 351 additions & 0 deletions integrations/agentbricks/docs/agent-tools.md

Large diffs are not rendered by default.

82 changes: 82 additions & 0 deletions integrations/agentbricks/docs/deploy-and-maintain.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
# Deploy and manage a generated agent

This guide is for owners of projects generated by `agentbricks init`. It explains which files and
dependency inputs control a deployment, how application resources and agent state behave across local
runs, redeployments, and deletion. For command syntax and options, see the [CLI command reference](../cli.md).
To upgrade a customized scaffold, see [Upgrade a generated agent project](upgrading-generated-projects.md).

## Deployment dependency inputs

The generated `pyproject.toml` declares Python packages; `app.yaml` runs `uv run start-server` in
Databricks Apps. Released packages can come from the configured package index (public PyPI by
default, or `agentbricks deploy --pip-index-url <index-url>`). For the advanced case of deploying an
unreleased package, pin a pushed Git commit that the Apps build can reach:

```toml
[tool.uv.sources]
databricks-agentbricks = { git = "https://github.com/<you>/databricks-ai-bridge", rev = "<pushed-sha>", subdirectory = "integrations/agentbricks" }
```

Commit and push the ref before deploying; the Apps build clones that commit. A local path or
`file://` source is unavailable inside the Apps build. Contributor-only local-path testing is
documented in [CONTRIBUTING.md](../CONTRIBUTING.md#testing-sdk--runtime-changes-in-a-scaffold).

`agentbricks deploy` uploads the project source but excludes the local `uv.lock`; the Apps build
resolves dependencies against its own index. The generated `>=` requirements can therefore resolve
to newer packages on a later deployment. Pin direct dependency versions in `pyproject.toml`, verify
the selected package index and reachable Git commits, then compare the local environment with the
deployed build logs. The current deploy flow does not provide a frozen transitive dependency graph
from the local lockfile. Keep `agent.toml` bindings and the chosen app name alongside the dependency
manifest so the same deployment targets the same managed resources.

## Resource and state lifecycle

The default managed-server template declares memory, session, and tracing names in `agent.toml`.
`init` writes those declarations without creating workspace resources. `deploy` resolves them in the
target workspace, creates missing resources, reuses accessible ones with matching names, creates or
updates the App, rolls out the source, then attempts the App's store and trace access grants. An existing
name that the caller cannot access causes an error. A grant failure can leave a deployed App without the
corresponding feature; inspect deploy warnings.
`agentbricks memory/sessions bind` and `unbind` edit `agent.toml`; they do not delete remote stores.
After a memory or session unbind, redeploy currently leaves any earlier `AGENT_MEMORY_STORE` or
`AGENT_SESSION_STORE` setting in `app.yaml` in place. Remove the stale setting from `app.yaml` before
redeploying if you want the App to stop using that store. A clean tracing unbind is removed on the
next deploy.
Default store names contain a six-letter token (`<name>-<token>-memory` and
`<name>-<token>-sessions`); use `agentbricks memory bind <name>` or
`agentbricks sessions bind <name>` to select existing stores. The default tracing experiment is
under `/Shared/agentbricks_traces/`; `agentbricks tracing list` shows available traces.

The Invocation Runtime Store row in the matrix applies only to projects with
`[agent].server = "agentbricks"`. A custom server defines its own protocol and does not receive a
managed Runtime Store.

| Resource or state | Created or reused | Local `dev` and restart | Redeploy and cleanup |
| --- | --- | --- | --- |
| Project files and dependencies | `init` copies a template; `dev` builds `.venv` from `pyproject.toml`. | Source files stay on disk. The local environment is reused until `dev --prepare-environment` rebuilds it. | Deploy syncs source and resolves dependencies again without the local `uv.lock`. App deletion leaves the local project alone. |
| Databricks App | Deploy creates the named App or reuses it, then updates compute and source. | `dev` serves the project locally without creating an App. | Redeploy with the same name updates that App; `agentbricks deployments delete <app-name>` deletes it. |
| Invocation Runtime Store | The current default deploy creates or reuses an App-owned database in the workspace's shared Lakebase project. An internal legacy path uses a per-App Lakebase project. | `dev` keeps invocation status, results, and events in process; they disappear on restart. | Deployed invocation records persist across restart and redeploy. Queued work can resume; active work needs a recovery handler and may run more than once. `agentbricks deployments delete` removes the managed store before the App; the legacy delete path does not explicitly remove its Lakebase project. |
| Managed tool access | Deploy reconciles direct App-auth tool grants from `agent.toml` before source upload, then finalizes Agent Bricks-owned App resources after rollout; request-user tools use the caller's permissions. | `dev` creates no App service principal or Apps grants. | Removing a tool binding removes Agent Bricks-owned Apps resources on redeploy. MCP and Workspace grants are additive; see [automatic App-identity access](agent-tools.md#automatic-app-identity-access-on-deploy). |
| Memory Store | Deploy creates a declared store if missing or reuses an accessible store by name, then attempts the App grant. | Managed long-term memory is off in `dev`. | Memory persists independently of the App; redeploy reuses the bound store. Unbinding or deleting the App does not delete it. Remove the stale `app.yaml` setting to detach it after unbind; use the separate store delete command when appropriate. |
| Session Store | Deploy creates or reuses a declared store by name, then attempts the App grant. | `dev` keeps conversation state in process, so a restart loses it. | A bound store preserves LangGraph checkpoints and OpenAI Agents SDK transcripts across restart and redeploy. OpenAI pending approval `RunState` stays in process. Unbinding or deleting the App does not delete the store; remove the stale `app.yaml` setting to detach it. |
| MLflow traces | `dev` uses a local MLflow server; deploy attempts to create or reuse the bound workspace experiment and grant App access. | Local traces are recorded in `.agentbricks/`; they remain on disk after `dev` stops. | Deployed traces remain in the workspace experiment. Unbinding removes the App's tracing configuration on a later clean deploy; it does not delete the experiment. |

Redeploying an older App can attach the current managed Runtime Store without migrating invocation
records from its legacy per-App Lakebase project. Managed-store cleanup errors retain the App for
retry; deleting the App directly bypasses that cleanup.

The Runtime Store tracks HTTP invocations, status, results, and event replay. The framework's
conversation history belongs to its Session Store when bound. The generated chat UI keeps its
session ID in browser local storage and sends that ID as the top-level `session_id` with each turn.
`DurableAgentServer` accepts an optional top-level `session_id` for generic handlers, but the generated
LangGraph and OpenAI Agents templates require a nonempty top-level `session_id` on every invocation.
API clients should reuse that value for conversation continuity and send a new invocation `id` for each
turn.
By default, the templates use the session ID as the state actor; request-user-authenticated
invocations namespace session state by user. LangGraph can resume from a matching checkpoint after
worker loss; the OpenAI Agents SDK template replays the input against its saved transcript.
Recovery can repeat external side effects, so make tools idempotent. For request-user-authenticated
tools, credentials are not persisted and background recovery is unsupported.

The default chat UI keeps its session ID across page reloads; see [invoke HTTP endpoints](../README.md#invoke-from-the-command-line)
for the explicit API request shape.
203 changes: 203 additions & 0 deletions integrations/agentbricks/docs/memory-and-sessions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,203 @@
# Memory and sessions

An agent needs two kinds of state: the state of the interaction it is handling right now, and the
durable knowledge it carries from one conversation to the next. Databricks provides a managed store
for each, both backed by Lakebase and usable from agents built on any framework:

- **Managed agent sessions** store an agent's session state for one interaction. Most commonly this
is the conversation history — the ordered transcript of messages, tool calls, and results — but it
can be any state a framework persists, such as a LangGraph graph. The agent reads it at the start
of a turn and appends to it as the interaction runs.
- **Managed agent memory** stores durable facts, preferences, and decisions that an agent recalls in
later, separate conversations through text search.

![Sessions and memory: the agent reads and appends one conversation's transcript in the session store, and recalls and saves durable facts in the memory store, which outlive any single conversation.](sessions_and_memory.png)

The examples below use the [`AgentKitClient` Python SDK](#agentkit-sdk). The same operations are
available as `agentbricks sessions` and `agentbricks memory` CLI commands; see the [CLI command
reference](../cli.md) for the complete command set.

## AgentKit SDK

Use `AgentKitClient` from `databricks_agentkit` with an authenticated Databricks
`WorkspaceClient`, or omit the client to use the default Databricks SDK authentication. Its
`memory_stores` and `session_stores` collections create, get, and list stores. A returned store
manages its entries or sessions; returned memories and sessions own their `update()` and
`delete()` operations.

All `list()` methods return iterators that consume server pages automatically. List `page_size` and
search `limit` values must be between 1 and 100. `session.list_items()` also auto-pages.

## Sessions

A **session store** holds **sessions**, and each session holds an ordered list of **session items**.
A session is one interaction — typically a conversation thread — grouped under an `actor_id` (who it
belongs to; set this from trusted application context, never a model- or user-supplied value) and
identified by a caller-chosen `session_id` (the service generates one if you omit it). Each item is
an opaque, JSON-compatible `data` value — a message, tool call, result, or reasoning block — that
Databricks stores and returns verbatim, in order, and never mutates once appended.

Create a store, start a session, append the conversation's turns, and read the history back on a later
request:

```python
from databricks.sdk import WorkspaceClient
from databricks_agentkit import AgentKitClient

agentkit = AgentKitClient(WorkspaceClient())

session_store = agentkit.session_stores.create("support-agent-sessions")
session = session_store.add(actor_id="customer-123", session_id="case-456")

session.append_items(
[
{"type": "message", "role": "user", "content": "I need help with my cluster."},
{"type": "message", "role": "assistant", "content": "Let's take a look."},
]
)

# On a later turn, reload the session and read its full history in order.
session = session_store.get("case-456")
history = [item.data for item in session.list_items()] # list_items auto-pages
```

A session can be **forked** into an independent branch: a new session seeded with the original's
history, linked back to its origin by `parent_session_id`. Fork the full history, or only up to a
specific item, to explore an alternate continuation without disturbing the original thread:

```python
branch = session.fork(actor_id="customer-123") # add up_to_item_id=... to branch up to one item
```

Deleting a session that has such descendants requires `session.delete(force=True)` to cascade.

## Memory

A **memory store** holds **memory entries**. Each entry is a free-form `content` string plus a short
`description` used for retrieval, keyed by three fields: `actor_id` (whose memory it is — set from
trusted application context, never a model- or user-supplied value), `path` (a filesystem-like key
within an actor, such as `/preferences/response-style.md`), and an optional `session_id` (the session
an entry came from, for provenance). An entry is uniquely identified by its `actor_id`, `path`, and
optional `session_id`.

Write an entry when the agent learns something durable, then recall it in a later, separate
conversation with a natural-language search. Results are ranked by full-text (BM25) relevance, up to
100 entries, with no pagination or vector similarity:

```python
from databricks.sdk import WorkspaceClient
from databricks_agentkit import AgentKitClient

agentkit = AgentKitClient(WorkspaceClient())

memory_store = agentkit.memory_stores.create("support-agent-memory")
memory_store.add(
actor_id="user-123",
path="/preferences/communication.md",
content="Prefers email over phone. Timezone: PST.",
description="User 123 communication preferences",
)

# In a later, separate conversation, recall what the agent knows about this user.
results = memory_store.search(actor_id="user-123", query="communication preferences", limit=10)
```

To browse rather than search, `memory_store.list(actor_id=..., path_prefix=...)` returns entries
directly.

> **`actor_id` partitions data; it is not access control.** Both stores are workspace-scoped and
> authorized at the store level, so any principal that can reach a store can read and write every
> actor's entries. For strict isolation between tenants or users, use a separate store per boundary.
> Grant another principal — such as your app's service principal — access with
> `session_store.grant_permission(principal_id)` or `memory_store.grant_permission(principal_id)`;
> `agentbricks deploy` attempts this grant for the deployed app.

## Framework adapters

In an agent configured with `server = "agentbricks"`, the framework adapter reads and appends
session state for you. The adapter resolves a bound store when one is configured. During local
`agentbricks dev`, sessions use process-local state and long-term memory is off; a deployed agent
gets the managed stores through the environment configured by `agentbricks deploy`.

### LangGraph

Pass `checkpointer()` when you build the agent and scope each run with `thread_config(session_id)`.
Pass the trusted actor as the second argument when the graph's state must be partitioned by actor.

```python
from databricks_agentkit.langgraph import checkpointer, thread_config

agent = create_agent(model=..., tools=[...], checkpointer=checkpointer())
result = await agent.ainvoke(inputs, config=thread_config(session_id))
```

Add the memory tools to the model and execution tool list. `memory_tools(actor)` exposes `remember`
and `recall` bound to one actor's partition. It resolves the store from the `[memory_store]` binding,
carried to the runtime by the `AGENT_MEMORY_STORE` environment variable that `agentbricks deploy`
injects. It returns no tools when no store is set, so the agent runs unchanged:

```python
from databricks_agentkit.langgraph import memory_tools

agent = create_agent(model=..., tools=[*your_tools, *memory_tools(actor)])
```

### OpenAI Agents SDK

The OpenAI Agents adapter exposes the same session and memory capabilities. Pass
`session_store(session_id)` to `Runner.run` and add `memory_tools(actor)` to the agent's tools. The
optional `actor` argument partitions durable session and memory data; derive it from trusted
application context.

```python
from agents import Agent, Runner
from databricks_agentkit.openai import memory_tools, session_store

agent = Agent(model=..., tools=[*your_tools, *memory_tools(actor)])
result = await Runner.run(
agent,
messages,
session=session_store(session_id, actor),
)
```

Both adapters resolve their store from the explicit argument first, then the corresponding
`AGENT_SESSION_STORE` or `AGENT_MEMORY_STORE` environment variable, and then the `agent.toml`
binding. With no managed store configured, the session helpers keep state in process and the memory
helper returns no tools. This is the local development behavior; deployment supplies the managed
store configuration.

## Declaring and provisioning stores

For a deployed agent, `agent.toml` declares which stores it uses and `agentbricks deploy` provisions
them — you do not create stores by hand for the deployed project. `agentbricks init` declares a
default memory and session store named from the project; override those names, point at stores you
already have, or let `deploy` create them:

To scaffold a project with default memory and session stores declared in `agent.toml`:

```sh
agentbricks init my-agent
```

To scaffold with specific store names instead:

```sh
agentbricks init my-agent --memory-store support-agent-memory --session-store support-agent-sessions
```

For an existing project, bind the stores in `agent.toml`, then deploy. Binding edits the manifest;
it does not create the remote stores:

```sh
# Run these commands from the project directory.
cd my-agent
agentbricks sessions bind support-agent-sessions
agentbricks memory bind support-agent-memory

# deploy creates any declared-but-missing store and attempts the App service-principal grant.
agentbricks deploy my-agent
```

Memory and session stores are independent resources: deleting one never affects the other. For runtime
hooks, endpoint behavior, and recovery, see the [runtime guide](../src/databricks_agentkit/runtime/README.md).
Loading
Loading