One line of changing text is now a list the reader can follow: Analyzing
the question → Searching the clinical library → Found 7 sources → Writing
the answer, plus any step the server reports (looking at an image, drawing
one, completing a cut-off reply). Steps are added when the stream says work
began and ticked when it says it ended, so the list is a record of the real
work, not an animation on a timer; nothing delays the answer.
The one hand-off the server cannot signal — retrieval runs before the stream
opens — is done by the stylesheet, not a JS timer, so the autosave debounce
keeps its clock to itself.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
The tail — the block still arriving — was shown as raw text until its blank
line came, which is the flash of asterisks, pipes and brackets the reader
saw on every paragraph. It is now parsed each frame like the finished part
(Open WebUI never shows the flash because it re-renders the whole message
per token; ours re-renders one block). Two exceptions stay text: a block
inside an unclosed code fence, so a half-written diagram is not handed to
its renderer every frame, and a table header with no body row yet.
JS, CSS and component HTML were cached for an hour, and the assistant's ES
module imports carry no version query, so every open browser kept the old
citation renderer for an hour after the deploy that replaced it. They now
revalidate on every load (no-cache with the ETag), which is a 304.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
The old renderer rewrote the text: it found "[n]" with regexes, renumbered
them, and swapped the result back in — which broke inside `arr[2][1]`, inside
HTML attributes, and whenever two turns disagreed about what "[3]" meant. It
also had a fallback markdown renderer of its own for when the rewrite
produced something markdown-it would not parse.
Now "[n]" is an inline rule registered on the same markdown-it instance that
renders everything else. The parser decides what is prose and what is code, a
link, or a URL, so the rule never sees "[1]" inside a code span, and it steps
aside for "[1](url)". Math is two more rules on the same parser instead of a
regex pre-pass, so "$" inside a URL is no longer math.
Identity vs display: the stored "[n]" and each card's id are the source's
identity (sourceNumber) and are never rewritten. The number a reader sees is
the order of first appearance, computed at render time from the token stream
(orderSourcesByCitation), so "one, then seven" cannot happen and a saved chat
re-opens pointing at the same cards it was saved with. Stored messages and
sources are untouched; export and the modal resolve by identity.
Translated HTML gets the same links through a TreeWalker over text nodes
(linkCitationsInHtml) rather than a regex over markup.
Deleted: renderCitationLinks, normalizeAdjacentCitationClusters, the
fallback renderer (fallbackMarkdown/renderMixedList/renderFallbackTable),
renderLatexText, CITATION_SCAN. Tests that asserted rewritten text now assert
token output; harnesses that render for real are given a parser.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
Measured against Open WebUI on the same question, sampling every 250ms:
it never showed a raw pipe. A real <table> appeared with 2 rows at
t+3.25s and grew to 4, then 6, as they arrived.
The blank-line rule alone cannot do that, and a table is the case it
serves worst: a markdown table contains no blank line, so the whole of
it stayed in the tail as plain text until the line *after* it landed,
then snapped into place. That is the markdown flash, on exactly the
content where it shows most.
So a table is now rendered while it is still arriving: once there is a
header, the |---| rule and one body row, the tail is rendered as
markdown rather than held as text, and every complete row that follows
joins it. A half-typed row is left out and appears a frame later, which
is what makes the table grow a row at a time.
Re-rendered each frame rather than appended, unlike a settled block: the
rows arriving next carry no header of their own, so they cannot be
parsed as a separate chunk. The tail is small, so the cost is small.
The status was the other half of that trace: Open WebUI keeps a skeleton
beside the streaming content until the answer is done. Ours removed the
status the instant the first token landed — the moment it becomes most
useful, because the answer is arriving *and* the assistant is still
working, drawing a figure or completing a cut-off reply. It now sits
above the partial answer, shimmering, until the final render replaces
the bubble.
And the wait itself says what is happening. Retrieval finishes before
the stream opens — deliberately, so a bad request still returns an error
rather than a stream — which makes that first line the only thing a
reader has during the slowest part. "Retrieving and synthesizing
references" described the software; "Searching the clinical library"
describes the work.
Mutation-tested: dropping the table rule, rendering the incomplete
trailing row, showing a table with no body row, or failing to restore
the status each frame all fail a test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
Markdown only means anything once a block is finished. Half a table is a
row of pipes; half a fence is a stray ```. Re-parsing the whole partial
answer every 180ms therefore flickered between a broken parse and the
real thing, and the previous answer to that was to abandon markdown
past 3500 characters or eight pipe rows and dump raw text into a <pre>
that had no CSS at all — inheriting the browser's black monospace
default. That is the dark block people saw while a table streamed.
Now the text is split at the last finished block — a blank line outside
a code fence — and everything before it is rendered once and *appended*.
Only the unfinished tail is plain text, styled as prose. Settled content
is never re-parsed and never rebuilt, so a diagram or chart that has
already drawn is not thrown away by the next token.
Two blank lines are not boundaries: the gap inside a loose list, which
would render one list as two each restarting at 1, and one inside
indented code. Telling the first from the perfectly good boundary
between a sentence and the list it introduces takes looking at both
sides of the gap, not just ahead.
renderEmbeddedBlocks now marks what it has drawn. mermaid.render is
async and replaces the element; a second Chart on one canvas throws.
Marked before the await, not after, or two frames race the same element.
Streaming is a preview: 'done' still renders the whole answer from
scratch, so anything transient here is settled by the final pass. That
is what makes appending safe.
Mutation-tested: removing fence tracking, the list guard, the backward
half of that guard, the append, the drawn marker, or the state reset in
fillMessageBubble each fail a test.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU