Commit graph

5 commits

Author SHA1 Message Date
Daniel
4f5687982d feat: Learning resources can be grounded in the clinical corpus
Some checks failed
Forgejo Android APK / Root app tests (push) Successful in 47s
Forgejo Docker Build / Root app tests (push) Successful in 46s
Forgejo Android APK / Build signed APK (push) Successful in 2m2s
Forgejo Docker Build / Build Docker image (push) Successful in 19s
Forgejo Docker Build / Deploy to the host (push) Failing after 1s
Learning generated everything from the model alone. A deck on bronchiolitis was
whatever the model remembered about bronchiolitis, with no connection to the
documents this institution actually indexed — while the assistant had been
searching that corpus all along.

Same collection, deliberately. mcp_bge_m3_1024 is already embedded with
openrouter-bge-m3 at 1024 dimensions; a second index over the same documents
with the same embedder would be a copy that drifts. What differs is the budget:
a chat answer wants a few tight excerpts because the reader is waiting, a
teaching resource synthesises a whole topic. So learning.search_limit and
learning.context_chars default to 30 and 2500 against the assistant's 8 and
1400, and are separate keys so tuning one cannot move the other.

Not unbounded, though. "No limit" only moves the ceiling from a setting to the
model's context window, where overflow truncates the middle of the prompt
silently — the worst place to lose source material. 60 results and 8000
characters per excerpt are the caps.

Opt in per generation: a resource on something the library does not cover is
better written without it than padded with the nearest unrelated excerpts.
Retrieval never fails a generation — the resource is then written from the model
alone, which is what happened before this existed — and every response reports
what it was grounded on, so a caller can say "24 excerpts" or "the library had
nothing on this" rather than quietly serving ungrounded material.

Verified against the live corpus: bronchiolitis, neonatal jaundice and febrile
seizure each returned 12 excerpts and ~23k characters from Nelson, Rudolph and
the Pediatric Clinical Practice Guidelines. A deck generated through the full
chain came back with textbook specificity that is not general recall —
bronchiolar diameter, birth-weight thresholds, the full pathogen list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-11 14:03:05 +02:00
Daniel
050a7d5241 feat: citation quality tracking, and the SSO settings fit a phone
Citation quality
- A citation naming a source that never came back is never rendered as a
  link, so it appears as plain text and nobody learns it happened. It is now
  measured on the server, where the answer and the sources both exist, so it
  is seen whether or not a browser rendered it.
- Four Prometheus counters feed a Grafana dashboard (Ped-AI Citation
  Quality): answers, citations written, answers affected, and individual
  unresolved markers. Only answers with at least one unresolved citation are
  stored, with the question and the titles retrieval returned, so an operator
  can judge whether retrieval came back thin or the model over-cited. Rows
  expire after 30 days: this is a quality signal, not a transcript log.
- Both answer paths are covered. /chat/stream is normal; /chat is the
  fallback the client uses when streaming fails, so auditing only the first
  would have hidden exactly the answers produced under failure.
- The tracker is resolved on demand and allowed to be absent. Seven test
  files load this route with a hand-built list of permitted imports, and
  adding a hard dependency would mean editing all seven — and the eighth
  written later would break. Observation must never be able to fail an
  answer, so a missing module simply means no tracking.
- Metric registration reuses an already-registered counter, because this
  module can legitimately load twice in one process.

SSO settings on mobile
- Six rows were laid out inline: flex with a 160px label and an input that
  would not shrink, so on a phone the row was wider than the screen with
  nothing to scroll and no way to reach the rest. They use .admin-row now,
  which already stacks below 640px. Verified at 390px and 360px: nothing
  off-screen, no sideways overflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dv6sqaY6Vq3ChZHMem3cnU
2026-09-10 23:44:43 +02:00
Daniel
09a4e02e39 feat: patient take home (copy/export/email), composer mic dictation, voice conversation mode, provider-aware translate languages 2026-09-08 20:01:46 +02:00
Daniel
8b072496e2 feat: Open WebUI-style assistant workspace — 3-column layout, markdown/math/code/tables, autosave with images, translation (LibreTranslate+DeepL), citation modal, Learning Hub moved in, handoff removed 2026-09-08 18:52:35 +02:00
Daniel
6be2d1375a feat: integrate durable image jobs/private assets into current core 2026-09-07 16:53:42 +02:00