Conversation
# Conflicts: # integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-langgraph/runtime/ui.py # integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-langgraph/ui/app.js # integrations/agentbricks/src/databricks_agentbricks/templates/ui/agent-openai/runtime/ui.py
This was referenced Oct 2, 2026
Contributor
Author
|
Superseded by #686, a much smaller independent PR targeting main (5 files, +9/-49). It fixes confirmed history/setup defects; broader local-state enhancements are deferred. The original branch and commits are preserved. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #681. The diff here contains only local-state work and its tested integration. Merge the first PR, then retarget this PR to
main.Local development currently ignores bound session/memory stores, while the binding commands tell users that restarting
devwill pick them up. The LangGraph chat UI also reads a browser-maintained transcript instead of the agent's checkpoint, so successful API invocations disappear from chat history. This change adds an explicit local workspace-store mode and makes UI history follow the agent's state and request identity.Fixes ML-70354.
Feedback addressed
Resulting behavior
dev, which disabled workspace memory and kept session history in-processdev --workspace-storesdev --workspace-storesgraph.aget_state()from the same checkpointer, including pending messages and approval interrupts; removes duplicate browser transcript writesPlain
devretains its existing local behavior. With--workspace-stores, history and long-term memory can survive restart, but local invocation status, replay events, background work, and OpenAI pending approval state remain in-process. Preflight verifies reads; writes are checked by the service when used. App-auth agents retain documented store-level sharing; this does not redefineactor_idas access control.Before/after validation
The same history/pagination regressions against baseline templates failed 5 tests: request-user OpenAI history, shared/request-user LangGraph managed history, and both frameworks' pagination. They pass after the fix. A real parallel LangGraph graph also verifies that a completed reply remains visible while another node pauses for approval, preserving the graph's public state semantics.
Live baseline deployment independently reproduced the LangGraph failure: foreground, streaming, background tool invocation, and a model-override invocation all completed, but
GET /api/demo/session/itemsreturned 0 items. Baseline OpenAI returned 8 items in the equivalent deployed probe. After this fix, the same live probe returned 12 LangGraph messages (including tool messages) and 8 OpenAI transcript items, with all four invocations completed. Both fixed Apps used combined source4389014570af6a6b14ca70bd0bbe51dd6066c772in ml-inference-staging.Live local testing used both newly generated framework projects, the selected staging workspace, a real
system.ai.claude-sonnet-4-5model, and uniquely named task-created Session/Memory Stores. Both generated agents used an editable source override to this checkout. The sequence was:/api/demo/sessions; submitted synthetic marker turns throughPOST /api/invocations, without browser transcript writes. Both returned HTTP 200/completed, andGET /api/demo/session/itemsincluded the marker.rememberwith a synthetic persistence fact. The real tool returnedstored; the Memory Store contained the entry.recall. The real tool output contained the previously saved fact for each framework (MEMORY_OPENAI_842/MEMORY_LANGGRAPH_842), proving memory retrieval rather than reuse of the prior conversation.NOT_FOUNDbefore server startup and printed the valid recovery commandagentbricks --profile <profile> sessions stores create --name ...; no store was created.NOT_FOUND. No task-created agent/MLflow processes remain.The first memory-list read encountered a transient platform
CANCELLEDerror; subsequent actual list/search/recall requests succeeded. Parallel local launches also collided on the Apps CLI's fixed default proxy port 8001, so successful CLI tests and restart checks were run sequentially. These environment issues were recorded separately from the successful state assertions.Checks
Generated OpenAI and LangGraph
tests/test_demo_ui.pyeach passed 24 tests. The follow-up recovery-command check passed 45 dev tests; the scaffold-helper loader follow-up passed the state regression file. No model/workspace mocks were used for the live local persistence assertions above. Unit tests replace remote transports explicitly.The final combined source passed 1,538 unit tests (14 skipped), 6 fresh-scaffold functional tests, 25 UI tests per framework, Ruff/format/type checks, and both frameworks' full browser and pause/resume checks. The final stack's application source matches the deployed/tested commit byte-for-byte; subsequent differences are test-only fixes.