Skip to content

[agentbricks] Fix chat rendering, invocation status, and model selection - #681

Open
shivam5 wants to merge 12 commits into
mainfrom
codex/bugbash-chat-models
Open

shivam5 wants to merge 12 commits into
mainfrom
codex/bugbash-chat-models

Conversation

@shivam5

@shivam5 shivam5 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Partially addresses ML-70352 and ML-70355.

Chat displayed Markdown literally, could report Ready after a failed stream, and stopped waiting after three minutes while a background agent was still running. The model picker hid services after its twentieth entry. Live Opus 5.5 reasoning tests also exposed LangGraph failures on absent token-usage details, duplicated final answers, and line breaks between token fragments.

This change renders assistant Markdown during streaming, on completion, and when loading history; Copy preserves the source Markdown. A pinned local markdown-it parser escapes raw HTML and disables images. Typed reasoning/signature blocks remain outside answer chat. Responses API fixes omit absent usage details, retain the last meaningful event across metadata events, and carry the output index on text deltas so LangChain assembles them correctly.

The UI uses Foreground / Background / Streaming, lists all discovered models, reports failed/disconnected streams, and waits for background completion without an arbitrary deadline. Polling backs off from 850 ms to a five-second cap to reduce traffic during long waits; this adds up to five seconds of completion-detection delay. A stuck run can still keep the tab busy. Closing/reloading stops browser polling without cancelling the agent; reconnect/stop-waiting UX is outside this patch.

Feedback addressed

Before / after

Case Before After
Markdown, both frameworks and all modes Literal headings, bold, backticks, and table syntax Rendered headings, bold, inline/fenced code, lists, links, and tables; original Markdown on Copy
Controlled history reload Plain text Same rendering as new replies
Live LangGraph + Opus Responses API Duplicated streaming answer; Foreground and Background fail when optional usage details are absent One answer; exact code preserved; all three modes complete
Failed or prematurely closed stream Can report Ready Error; composer re-enabled
185-second background invocation Three-minute timeout discards eventual result Active at 181 seconds, then one final answer; polling backs off to five seconds
Model list First 20 entries All discovered models; project default labeled

Validation

Tested source: f0208da. Generated UI assets were byte-compared with this commit; runs began before committing the identical backoff patch. Main baseline: ed3ed6b.

  • Live ml-inference-staging: newly generated OpenAI and LangGraph apps, Chrome → FastAPI → real adapters → Sonnet 4.5/Opus 5.5. Both frameworks passed all three modes, Markdown/exact code, one assistant answer, and codeword memory across model switching and a follow-up. Temporary test scaffolds enabled Responses reasoning.effort=medium and include=["reasoning.encrypted_content"] for Opus. LangGraph returned real encrypted payloads (820/900 characters) outside answer chat; OpenAI passed but did not expose encrypted items to the browser. The original visible-ciphertext symptom did not recur on current main; the recorded before/after demonstrates the reproduced duplicate-output/HTTP failures.
  • Controlled browser E2E: actual generated runtime/adapters plus local model fixtures, both frameworks. Markdown in all modes, split delimiters, exact clipboard, escaped HTML, disabled images, history hydration, failures, and model override with a real local tool roundtrip passed. A 185-second fixture made 41 polls over ~191 seconds and returned one answer. Separate browser response fixtures verify 404/503/network errors and premature SSE closure stop waiting and restore the composer.
  • Automated: LangChain chat-model suite 108 passed, 2 skipped; generated UI suites 24 passed per framework; JavaScript syntax and diff checks passed. Earlier usage/duplicate-output regression checks demonstrated failures before and passes after, including sync/async text assembly.

Live tests use local generated servers with the checkout's packages, including the changed LangChain integration; they do not deploy a Databricks App. Template updates apply to newly generated projects; existing projects need refreshed UI files. Discovery outside system.ai remains a follow-up.

@shivam5 shivam5 changed the title [agentbricks] Fix chat streaming, background feedback, and model selection [agentbricks] Fix chat status, streaming cadence, and model list truncation Oct 3, 2026

@jamesbxwu jamesbxwu left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you post any screenshots/videos to the new experience?

async function pollBackground(invocationId) {
const deadline = Date.now() + 180000;
while (Date.now() < deadline) {
while (true) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are there any downsides of polling forever vs just raising the timeout?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes: a stuck run can keep this tab busy and continue making status reads. Raising the timeout would bound the wait, but would still discard a healthy long-running result without cancelling the backend invocation.

In f0208da I kept terminal-state waiting and added a small backoff in both templates: 850 ms initially, increasing to a 5-second cap (about 12 reads/minute once capped). The tradeoff is up to five seconds before noticing completion, plus request latency. Requests remain sequential and the event list is capped at 80. Completion/failure, HTTP errors (including 404), and network errors stop polling. Closing/reloading the page stops client polling; it does not cancel the agent or automatically resume waiting after reload. Stop-waiting/reconnect support remains separate work.

Verified in recorded browser tests: a 185-second fixture remained active at 181 seconds and displayed its result once, with 41 polls over ~191 seconds. Both frameworks also stopped polling and restored the composer for 404, 503, and network failure; premature SSE closure reported an error.

@shivam5 shivam5 changed the title [agentbricks] Fix chat status, streaming cadence, and model list truncation [agentbricks] Fix chat rendering, invocation status, and model selection Oct 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants