Docs & cl

This commit is contained in:
Yiorgis Gozadinos 2026-04-20 10:00:34 +03:00
parent 68a9f191d2
commit fa87cf79c5
No known key found for this signature in database
10 changed files with 62 additions and 295 deletions

View file

@ -3,27 +3,38 @@
### Added
- **Document virtual filesystem in analysis sandbox**: Documents are mounted at `/documents/{id}/` with `metadata.json` (eager), `content.txt` (lazy), and `items.jsonl` (lazy). The agent uses standard Python `pathlib.Path` to browse and read document content and structure.
- **`doc_item_refs` and `labels` in search results**: Search results now include document item references and labels for cross-referencing with `items.jsonl`.
- **`--skill` flag for chat TUI**: `haiku-rag chat -s rag -s analysis` to enable specific skills. Defaults to `rag`. Use `-s analysis` for code execution, or both for the full toolset.
- **Document virtual filesystem in analysis sandbox**: Documents mounted at `/documents/{id}/` with `metadata.json` (eager), `content.txt` (lazy), and `items.jsonl` (lazy). Standard Python `pathlib.Path` for browsing and reading document content and structure.
- **`execute_code` skill tool**: Direct code execution in the sandbox, surfaced as individual AG-UI events in the chat TUI
- **`cite` skill tool**: Explicit citation registration with per-turn tracking via `citation_index` and `citations` fields in state
- **`--skill` flag for chat TUI**: `haiku-rag chat -s rag -s analysis` to enable specific skills
- **`--model` overrides all agents**: Chat, QA, research, and analysis agents all use the specified model
- **Collapsible program display in chat TUI**: Analysis code execution results shown as expandable code blocks
### Changed
- **Analysis sandbox `search()` now returns expanded results**: Search results automatically include surrounding context (adjacent paragraphs, complete tables, section content) via the document_items table
- **BREAKING**: Rename RLM agent to analysis agent throughout:
- **BREAKING: Flatten skill architecture**: Skill sub-agents now call `search`, `execute_code`, `cite`, `list_documents`, `get_document` directly — every tool call surfaces as an AG-UI event. Removes the 3rd agent layer where `ask`/`analyze`/`research` spawned inner agents whose tool calls were invisible.
- **BREAKING: Rename RLM agent to analysis agent** throughout:
- `agents/rlm/``agents/analysis/`, all classes renamed (`RLMResult` → `AnalysisResult`, etc.)
- `client.rlm()``client.analyze()`
- CLI: `haiku-rag rlm``haiku-rag analyze`
- MCP: `rlm_question``analyze`
- Config: `rlm:``analysis:` in YAML, `RLMConfig``AnalysisConfig`
- Skill: `rag-rlm``rag-analysis`, `skills/rlm.py``skills/analysis.py`
- State namespace: `"rlm"``"analysis"`
- Skill entrypoint: `rag-rlm``rag-analysis`
- **Analysis sandbox `search()` returns expanded results** with `doc_item_refs` and `labels` for cross-referencing with `items.jsonl`
- **`list_documents` skill tool** takes no parameters — returns all documents
- **Per-turn citation tracking**: `citation_index: dict[str, Citation]` (deduplicated) + `citations: list[list[str]]` (per-turn chunk IDs) replaces flat citation list
- **Search rate limiting**: Skill search tool enforces `config.qa.max_searches`
### Removed
- **`get_document()`, `get_docling_document()`**: Removed from analysis sandbox — replaced by the document virtual filesystem
- **`get_chunk()`**: Removed from analysis sandbox — search results now include expanded context automatically
- **`create_analysis_toolset()`**: Removed unused `tools/analysis.py` module.
- **`ask` skill tool**: Replaced by direct `search` + `cite` — the skill sub-agent searches and answers directly
- **`analyze` skill tool**: Replaced by direct `execute_code` + `search` + `cite`
- **`research` skill tool**: Removed from skill layer (still available via CLI `haiku-rag research` and MCP)
- **`get_document()`, `get_docling_document()`**: Removed from analysis sandbox — replaced by VFS
- **`get_chunk()`**: Removed from analysis sandbox — search results include expanded context
- **`create_analysis_toolset()`**: Removed unused `tools/analysis.py` module
- **`qa_history`, `reports` from skill state**: Conversational context handled by the outer chat agent
- **`combine_filters`, `build_document_filter`**: Removed from public API
## [0.40.1] - 2026-04-17

View file

@ -148,11 +148,11 @@ function ToolCallIndicator({
switch (toolName) {
case "search":
return <SearchIcon />;
case "ask":
return <MessageIcon />;
case "get_document":
return <FileIcon />;
case "execute_skill":
case "execute_code":
case "cite":
return <MessageIcon />;
default:
return <SearchIcon />;
@ -163,16 +163,16 @@ function ToolCallIndicator({
switch (toolName) {
case "search":
return "Search";
case "ask":
return "Ask";
case "get_document":
return "Document";
case "execute_skill":
return "Skill";
case "analyze":
return "Analyze";
case "research":
return "Research";
case "execute_code":
return "Code";
case "cite":
return "Cite";
case "list_documents":
return "Documents";
default:
return toolName;
}
@ -194,16 +194,16 @@ function ToolCallIndicator({
const query = args.query as string;
return <span className="tool-query">{query}</span>;
}
case "ask": {
const question = args.question as string;
return <span className="tool-query">{question}</span>;
}
case "get_document":
return <span className="tool-query">{args.query as string}</span>;
case "analyze":
return <span className="tool-query">{args.question as string}</span>;
case "research":
return <span className="tool-query">{args.question as string}</span>;
case "execute_code": {
const code = args.code as string | undefined;
return (
<span className="tool-query">
{code ? code.slice(0, 80) : "Running code..."}
</span>
);
}
default:
return <span>Processing...</span>;
}

View file

@ -52,7 +52,6 @@ Press `Ctrl+P` to open the command palette:
### Session Management
- Conversation history is maintained in memory for the session
- Previous Q/A pairs are automatically used as context for follow-up questions via the `ask` tool
- Citations are tracked per response and can be inspected
- Document filter restricts all searches to selected documents
- Clearing chat resets session state
@ -67,7 +66,7 @@ Browser-based conversational RAG with a CopilotKit frontend.
- Expandable citations with source documents, pages, and headings
- Visual grounding to view chunk source locations in documents
- Document filter to restrict searches to selected documents
- Session state view for inspecting accumulated Q&A history, citations, and documents
- Session state view for inspecting citations and search results
### Quick Start

View file

@ -1,233 +0,0 @@
# Architecture
High-level overview of haiku.rag components and data flow.
## System Overview
```mermaid
flowchart TB
subgraph Sources["Document Sources"]
Files[Files]
URLs[URLs]
Text[Text]
end
subgraph Processing["Processing Pipeline"]
Converter[Converter]
Chunker[Chunker]
Embedder[Embedder]
end
subgraph Storage["Storage Layer"]
LanceDB[(LanceDB)]
end
subgraph Agents["Agent Layer"]
QA[QA Agent]
Skill[RAG Skill]
Research[Research Graph]
Analysis[Analysis Agent]
end
subgraph Apps["Applications"]
CLI[CLI]
ChatTUI[Chat TUI]
WebApp[Web App]
Inspector[Inspector]
MCP[MCP Server]
end
Sources --> Converter
Converter --> Chunker
Chunker --> Embedder
Embedder --> LanceDB
LanceDB --> Agents
Agents --> Apps
```
## Core Components
### Storage Layer
LanceDB provides vector storage with full-text search capabilities:
- **DocumentRecord** - Document metadata and full content
- **ChunkRecord** - Text chunks with embeddings and structural metadata
- **SettingsRecord** - Database configuration and version info
Repositories handle CRUD operations:
- `DocumentRepository` - Create, read, update, delete documents
- `ChunkRepository` - Chunk management and hybrid search
- `SettingsRepository` - Configuration persistence
### Processing Pipeline
```mermaid
flowchart LR
Source[Source] --> Converter
Converter --> DoclingDoc[DoclingDocument]
DoclingDoc --> Chunker
Chunker --> Chunks[Chunks]
Chunks --> Embedder
Embedder --> Vectors[Vectors]
Vectors --> DB[(LanceDB)]
```
**Converters** transform sources into DoclingDocuments:
- `docling-local` - Local Docling processing
- `docling-serve` - Remote processing via docling-serve
**Chunkers** split documents into semantic chunks:
- Preserves document structure (tables, lists, code blocks)
- Maintains provenance (page numbers, headings)
- Configurable chunk size
**Embedders** generate vector representations:
| Provider | Models |
|----------|--------|
| Ollama | nomic-embed-text, mxbai-embed-large |
| OpenAI | text-embedding-3-small, text-embedding-3-large |
| VoyageAI | voyage-3, voyage-code-3 |
| vLLM | Any compatible model |
| LM Studio | Any compatible model |
### Agent Layer
Three agent types and a RAG skill for different use cases:
```mermaid
flowchart TB
subgraph QA["QA Agent"]
Q1[Question] --> S1[Search]
S1 --> A1[Answer]
end
subgraph Skill["RAG Skill"]
Q2[Question] --> Tools[Tool Selection]
Tools --> S2[Search / Ask / Analyze]
S2 --> A2[Answer]
A2 --> State[RAG State]
State -.-> Q2
end
subgraph Research["Research Graph"]
Q3[Question] --> Plan[Plan Next]
Plan --> SearchOne[Search One]
SearchOne --> Eval[Evaluate]
Eval -->|Continue| Plan
Eval -->|Done| Synthesize[Synthesize]
end
subgraph AnalysisAgent["Analysis Agent"]
Q4[Question] --> Code[Write Code]
Code --> Execute[Execute]
Execute --> Examine[Examine Results]
Examine -->|Iterate| Code
Examine -->|Done| A4[Answer]
end
```
**QA Agent** - Single-turn question answering:
- Searches for relevant chunks
- Expands context around results
- Generates answer with optional citations
**RAG Skill** - Multi-turn conversational RAG via [haiku.skills](https://github.com/ggozad/haiku.skills):
- Bundles search, list_documents, get_document, ask, analyze, and research tools
- Managed `RAGState` for session state (citations, QA history, document filters)
- Integrates with any pydantic-ai agent via `SkillToolset`
- Powers both the Chat TUI and web application
**Research Graph** - Iterative research workflow:
- Proposes one question at a time, evaluates the answer, then decides whether to continue
- Prior answers let the planner skip redundant searches
- Synthesizes structured report
**Analysis Agent** - Complex analytical tasks via code execution:
- Writes Python code to explore the knowledge base
- Executes in sandboxed environment
- Handles aggregation, computation, multi-document analysis
- Iterates until answer is found
### Applications
| Application | Interface | Use Case |
|-------------|-----------|----------|
| CLI | Command line | Scripts, one-off queries, batch processing |
| Chat TUI | Terminal | Interactive conversations |
| Web App | Browser | Team collaboration, visual interface |
| Inspector | Terminal | Database exploration, debugging |
| MCP Server | Protocol | AI assistant integration |
## Data Flow
### Document Ingestion
```mermaid
sequenceDiagram
participant User
participant CLI
participant Converter
participant Chunker
participant Embedder
participant DB as LanceDB
User->>CLI: add-src document.pdf
CLI->>Converter: Convert to DoclingDocument
Converter-->>CLI: DoclingDocument
CLI->>Chunker: Split into chunks
Chunker-->>CLI: Chunks with metadata
CLI->>Embedder: Generate embeddings
Embedder-->>CLI: Vectors
CLI->>DB: Store document + chunks
DB-->>User: Document ID
```
### Search and QA
```mermaid
sequenceDiagram
participant User
participant Agent
participant Embedder
participant DB as LanceDB
participant LLM
User->>Agent: Ask question
Agent->>Embedder: Embed query
Embedder-->>Agent: Query vector
Agent->>DB: Hybrid search
DB-->>Agent: Relevant chunks
Agent->>Agent: Expand context
Agent->>LLM: Generate answer
LLM-->>Agent: Answer + citations
Agent-->>User: Response
```
## Configuration
Configuration flows through the system:
```
CLI args → Environment variables → haiku.rag.yaml → Defaults
```
Key configuration areas:
- **Storage** - Database path, vacuum settings
- **Embeddings** - Provider, model, dimensions
- **Processing** - Chunk size, converter, chunker
- **Search** - Limits, context expansion
- **QA/Research** - Model, iterations, concurrency
- **Providers** - Ollama, vLLM, docling-serve URLs
See [Configuration](configuration/index.md) for details.

View file

@ -278,7 +278,7 @@ The generated package is a pip-installable Python package that registers as a `h
### Available Tools
`analyze`, `ask`, `get_document`, `list_documents`, `research`, `search`
`cite`, `execute_code`, `get_document`, `list_documents`, `search`
### Example
@ -287,7 +287,7 @@ The generated package is a pip-installable Python package that registers as a `h
haiku-rag create-skill \
--name medic \
--db /path/to/medic.lancedb \
--tools search,ask \
--tools search,cite \
--config-file /path/to/haiku.rag.yaml \
--description "Military medic knowledge base" \
--preamble "You are a military medic expert."

View file

@ -59,7 +59,6 @@ haiku-rag chat # Interactive conversation mode
- [Getting started](tutorial.md) - Tutorial
- [Installation](installation.md) - Install haiku.rag with different providers
- [Architecture](architecture.md) - System overview and data flow
- [Configuration](configuration/index.md) - Environment variables and settings
- [CLI](cli.md) - Command line interface usage
- [Python](python.md) - Python API reference

View file

@ -19,13 +19,10 @@ skill = create_skill(db_path=db_path, config=config)
| Tool | Purpose |
|------|---------|
| `analyze(question, document?, filter?)` | Answer analytical questions using code execution |
**Parameters:**
- `question` — The analytical question to answer.
- `document` — Optional document ID or title to pre-load for analysis.
- `filter` — Optional SQL WHERE clause to filter documents.
| `search(query, limit?)` | Hybrid search (vector + full-text) with context expansion |
| `list_documents()` | List all documents in the knowledge base |
| `execute_code(code)` | Execute Python code in a sandboxed interpreter with VFS access |
| `cite(chunk_ids)` | Register chunk IDs as citations for the current answer |
## State
@ -34,16 +31,16 @@ The skill manages an `AnalysisState` under the `"analysis"` namespace:
```python
class AnalysisState(BaseModel):
document_filter: str | None = None
analyses: list[AnalysisEntry] = []
class AnalysisEntry(BaseModel):
question: str
answer: str
program: str | None = None
executions: list[CodeExecutionEntry] = []
citation_index: dict[str, Citation] = {}
citations: list[list[str]] = []
searches: dict[str, list[SearchResult]] = {}
```
- **document_filter** — SQL WHERE clause applied to `analyze` calls (combined with any explicit `filter` parameter). Set this to scope analysis to specific documents.
- **analyses** — Each `analyze` call appends an `AnalysisEntry` with the question, answer, and executed program.
- **document_filter** — SQL WHERE clause applied to `search` and `list_documents` calls.
- **executions** — Each `execute_code` call appends a `CodeExecutionEntry` with code, stdout, stderr, and success status.
- **citation_index** / **citations** — Same per-turn citation tracking as the RAG skill.
- **searches** — Search results from both the `search` tool and sandbox-internal searches.
## Usage with RAG Skill
@ -67,4 +64,4 @@ agent = Agent(
)
```
See the [Analysis Agent](../agents/analysis.md) documentation for details on how the underlying agent works.
See the [Analysis Agent](../agents/analysis.md) documentation for details on how the underlying sandbox works.

View file

@ -47,7 +47,7 @@ Use `create-skill` to generate a standalone skill package with an embedded datab
haiku-rag create-skill \
--name recipes \
--db /path/to/recipes.lancedb \
--tools search,ask \
--tools search,cite \
--description "Recipe knowledge base" \
--preamble "You are a recipe expert."
```

View file

@ -1,6 +1,6 @@
# RAG Skill
The RAG skill is the primary way to use haiku.rag tools. It bundles search, Q&A, document browsing, and research into a single skill with managed state.
The RAG skill is the primary way to use haiku.rag tools. It bundles search, document browsing, and citation management into a single skill with managed state.
## `create_skill(db_path?, config?)`
@ -20,10 +20,9 @@ skill = create_skill(db_path=db_path, config=config)
| Tool | Purpose |
|------|---------|
| `search(query, limit?)` | Hybrid search (vector + full-text) with context expansion |
| `list_documents(limit?, offset?, filter?)` | Paginated document listing |
| `list_documents()` | List all documents in the knowledge base |
| `get_document(query)` | Retrieve a document by ID, title, or URI |
| `ask(question)` | Q&A with citations via the QA agent |
| `research(question)` | Deep multi-agent research producing comprehensive reports |
| `cite(chunk_ids)` | Register chunk IDs as citations for the current answer |
## State
@ -31,17 +30,13 @@ The skill manages a `RAGState` under the `"rag"` namespace:
```python
class RAGState(BaseModel):
citations: list[Citation] = []
qa_history: list[QAHistoryEntry] = []
citation_index: dict[str, Citation] = {}
citations: list[list[str]] = []
document_filter: str | None = None
searches: dict[str, list[SearchResult]] = {}
documents: list[DocumentInfo] = []
reports: list[ResearchEntry] = []
```
- **citations** — Accumulated citations from `ask` calls, with sequential indexing across calls.
- **qa_history** — Questions and answers from `ask` calls. Prior Q&A is used as context for follow-up questions when embeddings are similar.
- **document_filter** — SQL WHERE clause applied to `search`, `list_documents`, `ask`, and `research` calls. Set this to scope queries to specific documents.
- **citation_index** — All citations indexed by chunk ID (deduplicated across turns).
- **citations** — Per-turn lists of chunk IDs registered via the `cite` tool.
- **document_filter** — SQL WHERE clause applied to `search` and `list_documents` calls. Set this to scope queries to specific documents.
- **searches** — Search results keyed by query string.
- **documents** — Documents seen via `list_documents` or `get_document` (deduplicated by ID).
- **reports** — Research reports from `research` calls.

View file

@ -59,7 +59,6 @@ nav:
- index.md
- Getting started: tutorial.md
- Installation: installation.md
- Architecture: architecture.md
- Configuration:
- configuration/index.md
- Providers: configuration/providers.md