Skip to content

[agentbricks] Enable local workspace stores and authoritative session history - #682

Closed
shivam5 wants to merge 5 commits into
codex/bugbash-chat-modelsfrom
codex/bugbash-local-state
Closed

shivam5 wants to merge 5 commits into
codex/bugbash-chat-modelsfrom
codex/bugbash-local-state

Conversation

@shivam5

@shivam5 shivam5 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #681. The diff here contains only local-state work and its tested integration. Merge the first PR, then retarget this PR to main.

Local development currently ignores bound session/memory stores, while the binding commands tell users that restarting dev will pick them up. The LangGraph chat UI also reads a browser-maintained transcript instead of the agent's checkpoint, so successful API invocations disappear from chat history. This change adds an explicit local workspace-store mode and makes UI history follow the agent's state and request identity.

Fixes ML-70354.

Feedback addressed

  • Chang Shi — Agentbricks CLI Feedback: Functionality → Sessions, Steps 1–3, Functionality → Memory, and Deployment. Store creation, project binding, and conversation creation were confusing; local state appeared ineffective even when deployment worked.
  • Fabian — Production Grade Document Chatbot: Codex friction points, items 1 and 2 — Local-to-production behavioral parity; A coherent identity, session, and history model.

Resulting behavior

Trigger Before After
Bind an existing store and run locally Bind output recommended plain dev, which disabled workspace memory and kept session history in-process Bind output distinguishes saved configuration from verified readiness and points to dev --workspace-stores
dev --workspace-stores Unsupported Resolves the current project's bindings in the selected profile, verifies read access, and injects the resolved store IDs without creating resources or grants
Bound store missing/inaccessible No way to validate/use it in local dev Startup stops before launching app/tracing processes with store-specific recovery guidance
LangGraph turn submitted through the API, then history loaded in UI Managed history was empty unless the browser separately wrote messages Reads graph.aget_state() from the same checkpointer, including pending messages and approval interrupts; removes duplicate browser transcript writes
Request-user-auth agent history/memory routes UI used public IDs while the invocation runtime used namespaced IDs Uses the same runtime identity mapping; public browser IDs remain stable; current API-created sessions can be listed without UI metadata
Session with more than one page of items UI read only the first page Retrieves all pages in write order

Plain dev retains its existing local behavior. With --workspace-stores, history and long-term memory can survive restart, but local invocation status, replay events, background work, and OpenAI pending approval state remain in-process. Preflight verifies reads; writes are checked by the service when used. App-auth agents retain documented store-level sharing; this does not redefine actor_id as access control.

Before/after validation

The same history/pagination regressions against baseline templates failed 5 tests: request-user OpenAI history, shared/request-user LangGraph managed history, and both frameworks' pagination. They pass after the fix. A real parallel LangGraph graph also verifies that a completed reply remains visible while another node pauses for approval, preserving the graph's public state semantics.

Live baseline deployment independently reproduced the LangGraph failure: foreground, streaming, background tool invocation, and a model-override invocation all completed, but GET /api/demo/session/items returned 0 items. Baseline OpenAI returned 8 items in the equivalent deployed probe. After this fix, the same live probe returned 12 LangGraph messages (including tool messages) and 8 OpenAI transcript items, with all four invocations completed. Both fixed Apps used combined source 4389014570af6a6b14ca70bd0bbe51dd6066c772 in ml-inference-staging.

Live local testing used both newly generated framework projects, the selected staging workspace, a real system.ai.claude-sonnet-4-5 model, and uniquely named task-created Session/Memory Stores. Both generated agents used an editable source override to this checkout. The sequence was:

agentbricks --profile <profile> sessions stores create --name <test-session-store>
agentbricks --profile <profile> memory stores create --name <test-memory-store>
agentbricks --profile <profile> init --framework <openai-or-langgraph> <test-project>
cd <test-project>
agentbricks sessions bind <test-session-store>
agentbricks memory bind <test-memory-store>
# pyproject.toml: [tool.uv.sources] points databricks-agentbricks and databricks-ai-bridge
# to the tested checkout, as documented in CONTRIBUTING.md.
uv sync
agentbricks --profile <profile> dev --workspace-stores --no-prepare-environment --app-port <port>
  1. Created a conversation through /api/demo/sessions; submitted synthetic marker turns through POST /api/invocations, without browser transcript writes. Both returned HTTP 200/completed, and GET /api/demo/session/items included the marker.
  2. Asked each agent to call remember with a synthetic persistence fact. The real tool returned stored; the Memory Store contained the entry.
  3. Stopped the entire local process and restarted with the same command. The existing conversation was still visible: OpenAI 4 transcript items; LangGraph 8 checkpointed messages.
  4. Submitted a new conversation with the same actor and asked the agent to use recall. The real tool output contained the previously saved fact for each framework (MEMORY_OPENAI_842 / MEMORY_LANGGRAPH_842), proving memory retrieval rather than reuse of the prior conversation.
  5. Ran the command with a nonexistent bound session store. It exited 1 with NOT_FOUND before server startup and printed the valid recovery command agentbricks --profile <profile> sessions stores create --name ...; no store was created.
  6. Deleted only the two task-created stores and verified both return NOT_FOUND. No task-created agent/MLflow processes remain.

The first memory-list read encountered a transient platform CANCELLED error; subsequent actual list/search/recall requests succeeded. Parallel local launches also collided on the Apps CLI's fixed default proxy port 8001, so successful CLI tests and restart checks were run sequentially. These environment issues were recorded separately from the successful state assertions.

Checks

cd integrations/agentbricks
uv run pytest tests/unit_tests -q
# 1505 passed, 14 skipped
uv run pytest tests/unit_tests/template_state_history_test.py tests/unit_tests/dev_test.py \
  tests/unit_tests/bind_cli_test.py tests/unit_tests/sessions_cli_test.py \
  tests/unit_tests/memory_cli_test.py -q
# 79 passed, 2 skipped
uv run ty check
uv run ruff check

Generated OpenAI and LangGraph tests/test_demo_ui.py each passed 24 tests. The follow-up recovery-command check passed 45 dev tests; the scaffold-helper loader follow-up passed the state regression file. No model/workspace mocks were used for the live local persistence assertions above. Unit tests replace remote transports explicitly.

The final combined source passed 1,538 unit tests (14 skipped), 6 fresh-scaffold functional tests, 25 UI tests per framework, Ruff/format/type checks, and both frameworks' full browser and pause/resume checks. The final stack's application source matches the deployed/tested commit byte-for-byte; subsequent differences are test-only fixes.

# Conflicts:
#	integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-langgraph/runtime/ui.py
#	integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-langgraph/ui/app.js
#	integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-openai/runtime/ui.py
@shivam5

shivam5 commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #686, a much smaller independent PR targeting main (5 files, +9/-49). It fixes confirmed history/setup defects; broader local-state enhancements are deferred. The original branch and commits are preserved.

@shivam5 shivam5 closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant