Docs & cl
This commit is contained in:
parent
68a9f191d2
commit
fa87cf79c5
10 changed files with 62 additions and 295 deletions
31
CHANGELOG.md
31
CHANGELOG.md
|
|
@ -3,27 +3,38 @@
|
|||
|
||||
### Added
|
||||
|
||||
- **Document virtual filesystem in analysis sandbox**: Documents are mounted at `/documents/{id}/` with `metadata.json` (eager), `content.txt` (lazy), and `items.jsonl` (lazy). The agent uses standard Python `pathlib.Path` to browse and read document content and structure.
|
||||
- **`doc_item_refs` and `labels` in search results**: Search results now include document item references and labels for cross-referencing with `items.jsonl`.
|
||||
- **`--skill` flag for chat TUI**: `haiku-rag chat -s rag -s analysis` to enable specific skills. Defaults to `rag`. Use `-s analysis` for code execution, or both for the full toolset.
|
||||
- **Document virtual filesystem in analysis sandbox**: Documents mounted at `/documents/{id}/` with `metadata.json` (eager), `content.txt` (lazy), and `items.jsonl` (lazy). Standard Python `pathlib.Path` for browsing and reading document content and structure.
|
||||
- **`execute_code` skill tool**: Direct code execution in the sandbox, surfaced as individual AG-UI events in the chat TUI
|
||||
- **`cite` skill tool**: Explicit citation registration with per-turn tracking via `citation_index` and `citations` fields in state
|
||||
- **`--skill` flag for chat TUI**: `haiku-rag chat -s rag -s analysis` to enable specific skills
|
||||
- **`--model` overrides all agents**: Chat, QA, research, and analysis agents all use the specified model
|
||||
- **Collapsible program display in chat TUI**: Analysis code execution results shown as expandable code blocks
|
||||
|
||||
### Changed
|
||||
|
||||
- **Analysis sandbox `search()` now returns expanded results**: Search results automatically include surrounding context (adjacent paragraphs, complete tables, section content) via the document_items table
|
||||
- **BREAKING**: Rename RLM agent to analysis agent throughout:
|
||||
- **BREAKING: Flatten skill architecture**: Skill sub-agents now call `search`, `execute_code`, `cite`, `list_documents`, `get_document` directly — every tool call surfaces as an AG-UI event. Removes the 3rd agent layer where `ask`/`analyze`/`research` spawned inner agents whose tool calls were invisible.
|
||||
- **BREAKING: Rename RLM agent to analysis agent** throughout:
|
||||
- `agents/rlm/` → `agents/analysis/`, all classes renamed (`RLMResult` → `AnalysisResult`, etc.)
|
||||
- `client.rlm()` → `client.analyze()`
|
||||
- CLI: `haiku-rag rlm` → `haiku-rag analyze`
|
||||
- MCP: `rlm_question` → `analyze`
|
||||
- Config: `rlm:` → `analysis:` in YAML, `RLMConfig` → `AnalysisConfig`
|
||||
- Skill: `rag-rlm` → `rag-analysis`, `skills/rlm.py` → `skills/analysis.py`
|
||||
- State namespace: `"rlm"` → `"analysis"`
|
||||
- Skill entrypoint: `rag-rlm` → `rag-analysis`
|
||||
- **Analysis sandbox `search()` returns expanded results** with `doc_item_refs` and `labels` for cross-referencing with `items.jsonl`
|
||||
- **`list_documents` skill tool** takes no parameters — returns all documents
|
||||
- **Per-turn citation tracking**: `citation_index: dict[str, Citation]` (deduplicated) + `citations: list[list[str]]` (per-turn chunk IDs) replaces flat citation list
|
||||
- **Search rate limiting**: Skill search tool enforces `config.qa.max_searches`
|
||||
|
||||
### Removed
|
||||
|
||||
- **`get_document()`, `get_docling_document()`**: Removed from analysis sandbox — replaced by the document virtual filesystem
|
||||
- **`get_chunk()`**: Removed from analysis sandbox — search results now include expanded context automatically
|
||||
- **`create_analysis_toolset()`**: Removed unused `tools/analysis.py` module.
|
||||
- **`ask` skill tool**: Replaced by direct `search` + `cite` — the skill sub-agent searches and answers directly
|
||||
- **`analyze` skill tool**: Replaced by direct `execute_code` + `search` + `cite`
|
||||
- **`research` skill tool**: Removed from skill layer (still available via CLI `haiku-rag research` and MCP)
|
||||
- **`get_document()`, `get_docling_document()`**: Removed from analysis sandbox — replaced by VFS
|
||||
- **`get_chunk()`**: Removed from analysis sandbox — search results include expanded context
|
||||
- **`create_analysis_toolset()`**: Removed unused `tools/analysis.py` module
|
||||
- **`qa_history`, `reports` from skill state**: Conversational context handled by the outer chat agent
|
||||
- **`combine_filters`, `build_document_filter`**: Removed from public API
|
||||
|
||||
## [0.40.1] - 2026-04-17
|
||||
|
||||
|
|
|
|||
|
|
@ -148,11 +148,11 @@ function ToolCallIndicator({
|
|||
switch (toolName) {
|
||||
case "search":
|
||||
return <SearchIcon />;
|
||||
case "ask":
|
||||
return <MessageIcon />;
|
||||
case "get_document":
|
||||
return <FileIcon />;
|
||||
case "execute_skill":
|
||||
case "execute_code":
|
||||
case "cite":
|
||||
return <MessageIcon />;
|
||||
default:
|
||||
return <SearchIcon />;
|
||||
|
|
@ -163,16 +163,16 @@ function ToolCallIndicator({
|
|||
switch (toolName) {
|
||||
case "search":
|
||||
return "Search";
|
||||
case "ask":
|
||||
return "Ask";
|
||||
case "get_document":
|
||||
return "Document";
|
||||
case "execute_skill":
|
||||
return "Skill";
|
||||
case "analyze":
|
||||
return "Analyze";
|
||||
case "research":
|
||||
return "Research";
|
||||
case "execute_code":
|
||||
return "Code";
|
||||
case "cite":
|
||||
return "Cite";
|
||||
case "list_documents":
|
||||
return "Documents";
|
||||
default:
|
||||
return toolName;
|
||||
}
|
||||
|
|
@ -194,16 +194,16 @@ function ToolCallIndicator({
|
|||
const query = args.query as string;
|
||||
return <span className="tool-query">{query}</span>;
|
||||
}
|
||||
case "ask": {
|
||||
const question = args.question as string;
|
||||
return <span className="tool-query">{question}</span>;
|
||||
}
|
||||
case "get_document":
|
||||
return <span className="tool-query">{args.query as string}</span>;
|
||||
case "analyze":
|
||||
return <span className="tool-query">{args.question as string}</span>;
|
||||
case "research":
|
||||
return <span className="tool-query">{args.question as string}</span>;
|
||||
case "execute_code": {
|
||||
const code = args.code as string | undefined;
|
||||
return (
|
||||
<span className="tool-query">
|
||||
{code ? code.slice(0, 80) : "Running code..."}
|
||||
</span>
|
||||
);
|
||||
}
|
||||
default:
|
||||
return <span>Processing...</span>;
|
||||
}
|
||||
|
|
|
|||
|
|
@ -52,7 +52,6 @@ Press `Ctrl+P` to open the command palette:
|
|||
### Session Management
|
||||
|
||||
- Conversation history is maintained in memory for the session
|
||||
- Previous Q/A pairs are automatically used as context for follow-up questions via the `ask` tool
|
||||
- Citations are tracked per response and can be inspected
|
||||
- Document filter restricts all searches to selected documents
|
||||
- Clearing chat resets session state
|
||||
|
|
@ -67,7 +66,7 @@ Browser-based conversational RAG with a CopilotKit frontend.
|
|||
- Expandable citations with source documents, pages, and headings
|
||||
- Visual grounding to view chunk source locations in documents
|
||||
- Document filter to restrict searches to selected documents
|
||||
- Session state view for inspecting accumulated Q&A history, citations, and documents
|
||||
- Session state view for inspecting citations and search results
|
||||
|
||||
### Quick Start
|
||||
|
||||
|
|
|
|||
|
|
@ -1,233 +0,0 @@
|
|||
# Architecture
|
||||
|
||||
High-level overview of haiku.rag components and data flow.
|
||||
|
||||
## System Overview
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Sources["Document Sources"]
|
||||
Files[Files]
|
||||
URLs[URLs]
|
||||
Text[Text]
|
||||
end
|
||||
|
||||
subgraph Processing["Processing Pipeline"]
|
||||
Converter[Converter]
|
||||
Chunker[Chunker]
|
||||
Embedder[Embedder]
|
||||
end
|
||||
|
||||
subgraph Storage["Storage Layer"]
|
||||
LanceDB[(LanceDB)]
|
||||
end
|
||||
|
||||
subgraph Agents["Agent Layer"]
|
||||
QA[QA Agent]
|
||||
Skill[RAG Skill]
|
||||
Research[Research Graph]
|
||||
Analysis[Analysis Agent]
|
||||
end
|
||||
|
||||
subgraph Apps["Applications"]
|
||||
CLI[CLI]
|
||||
ChatTUI[Chat TUI]
|
||||
WebApp[Web App]
|
||||
Inspector[Inspector]
|
||||
MCP[MCP Server]
|
||||
end
|
||||
|
||||
Sources --> Converter
|
||||
Converter --> Chunker
|
||||
Chunker --> Embedder
|
||||
Embedder --> LanceDB
|
||||
|
||||
LanceDB --> Agents
|
||||
Agents --> Apps
|
||||
```
|
||||
|
||||
## Core Components
|
||||
|
||||
### Storage Layer
|
||||
|
||||
LanceDB provides vector storage with full-text search capabilities:
|
||||
|
||||
- **DocumentRecord** - Document metadata and full content
|
||||
- **ChunkRecord** - Text chunks with embeddings and structural metadata
|
||||
- **SettingsRecord** - Database configuration and version info
|
||||
|
||||
Repositories handle CRUD operations:
|
||||
|
||||
- `DocumentRepository` - Create, read, update, delete documents
|
||||
- `ChunkRepository` - Chunk management and hybrid search
|
||||
- `SettingsRepository` - Configuration persistence
|
||||
|
||||
### Processing Pipeline
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Source[Source] --> Converter
|
||||
Converter --> DoclingDoc[DoclingDocument]
|
||||
DoclingDoc --> Chunker
|
||||
Chunker --> Chunks[Chunks]
|
||||
Chunks --> Embedder
|
||||
Embedder --> Vectors[Vectors]
|
||||
Vectors --> DB[(LanceDB)]
|
||||
```
|
||||
|
||||
**Converters** transform sources into DoclingDocuments:
|
||||
|
||||
- `docling-local` - Local Docling processing
|
||||
- `docling-serve` - Remote processing via docling-serve
|
||||
|
||||
**Chunkers** split documents into semantic chunks:
|
||||
|
||||
- Preserves document structure (tables, lists, code blocks)
|
||||
- Maintains provenance (page numbers, headings)
|
||||
- Configurable chunk size
|
||||
|
||||
**Embedders** generate vector representations:
|
||||
|
||||
| Provider | Models |
|
||||
|----------|--------|
|
||||
| Ollama | nomic-embed-text, mxbai-embed-large |
|
||||
| OpenAI | text-embedding-3-small, text-embedding-3-large |
|
||||
| VoyageAI | voyage-3, voyage-code-3 |
|
||||
| vLLM | Any compatible model |
|
||||
| LM Studio | Any compatible model |
|
||||
|
||||
### Agent Layer
|
||||
|
||||
Three agent types and a RAG skill for different use cases:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph QA["QA Agent"]
|
||||
Q1[Question] --> S1[Search]
|
||||
S1 --> A1[Answer]
|
||||
end
|
||||
|
||||
subgraph Skill["RAG Skill"]
|
||||
Q2[Question] --> Tools[Tool Selection]
|
||||
Tools --> S2[Search / Ask / Analyze]
|
||||
S2 --> A2[Answer]
|
||||
A2 --> State[RAG State]
|
||||
State -.-> Q2
|
||||
end
|
||||
|
||||
subgraph Research["Research Graph"]
|
||||
Q3[Question] --> Plan[Plan Next]
|
||||
Plan --> SearchOne[Search One]
|
||||
SearchOne --> Eval[Evaluate]
|
||||
Eval -->|Continue| Plan
|
||||
Eval -->|Done| Synthesize[Synthesize]
|
||||
end
|
||||
|
||||
subgraph AnalysisAgent["Analysis Agent"]
|
||||
Q4[Question] --> Code[Write Code]
|
||||
Code --> Execute[Execute]
|
||||
Execute --> Examine[Examine Results]
|
||||
Examine -->|Iterate| Code
|
||||
Examine -->|Done| A4[Answer]
|
||||
end
|
||||
```
|
||||
|
||||
**QA Agent** - Single-turn question answering:
|
||||
|
||||
- Searches for relevant chunks
|
||||
- Expands context around results
|
||||
- Generates answer with optional citations
|
||||
|
||||
**RAG Skill** - Multi-turn conversational RAG via [haiku.skills](https://github.com/ggozad/haiku.skills):
|
||||
|
||||
- Bundles search, list_documents, get_document, ask, analyze, and research tools
|
||||
- Managed `RAGState` for session state (citations, QA history, document filters)
|
||||
- Integrates with any pydantic-ai agent via `SkillToolset`
|
||||
- Powers both the Chat TUI and web application
|
||||
|
||||
**Research Graph** - Iterative research workflow:
|
||||
|
||||
- Proposes one question at a time, evaluates the answer, then decides whether to continue
|
||||
- Prior answers let the planner skip redundant searches
|
||||
- Synthesizes structured report
|
||||
|
||||
**Analysis Agent** - Complex analytical tasks via code execution:
|
||||
|
||||
- Writes Python code to explore the knowledge base
|
||||
- Executes in sandboxed environment
|
||||
- Handles aggregation, computation, multi-document analysis
|
||||
- Iterates until answer is found
|
||||
|
||||
### Applications
|
||||
|
||||
| Application | Interface | Use Case |
|
||||
|-------------|-----------|----------|
|
||||
| CLI | Command line | Scripts, one-off queries, batch processing |
|
||||
| Chat TUI | Terminal | Interactive conversations |
|
||||
| Web App | Browser | Team collaboration, visual interface |
|
||||
| Inspector | Terminal | Database exploration, debugging |
|
||||
| MCP Server | Protocol | AI assistant integration |
|
||||
|
||||
## Data Flow
|
||||
|
||||
### Document Ingestion
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant User
|
||||
participant CLI
|
||||
participant Converter
|
||||
participant Chunker
|
||||
participant Embedder
|
||||
participant DB as LanceDB
|
||||
|
||||
User->>CLI: add-src document.pdf
|
||||
CLI->>Converter: Convert to DoclingDocument
|
||||
Converter-->>CLI: DoclingDocument
|
||||
CLI->>Chunker: Split into chunks
|
||||
Chunker-->>CLI: Chunks with metadata
|
||||
CLI->>Embedder: Generate embeddings
|
||||
Embedder-->>CLI: Vectors
|
||||
CLI->>DB: Store document + chunks
|
||||
DB-->>User: Document ID
|
||||
```
|
||||
|
||||
### Search and QA
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant User
|
||||
participant Agent
|
||||
participant Embedder
|
||||
participant DB as LanceDB
|
||||
participant LLM
|
||||
|
||||
User->>Agent: Ask question
|
||||
Agent->>Embedder: Embed query
|
||||
Embedder-->>Agent: Query vector
|
||||
Agent->>DB: Hybrid search
|
||||
DB-->>Agent: Relevant chunks
|
||||
Agent->>Agent: Expand context
|
||||
Agent->>LLM: Generate answer
|
||||
LLM-->>Agent: Answer + citations
|
||||
Agent-->>User: Response
|
||||
```
|
||||
|
||||
## Configuration
|
||||
|
||||
Configuration flows through the system:
|
||||
|
||||
```
|
||||
CLI args → Environment variables → haiku.rag.yaml → Defaults
|
||||
```
|
||||
|
||||
Key configuration areas:
|
||||
|
||||
- **Storage** - Database path, vacuum settings
|
||||
- **Embeddings** - Provider, model, dimensions
|
||||
- **Processing** - Chunk size, converter, chunker
|
||||
- **Search** - Limits, context expansion
|
||||
- **QA/Research** - Model, iterations, concurrency
|
||||
- **Providers** - Ollama, vLLM, docling-serve URLs
|
||||
|
||||
See [Configuration](configuration/index.md) for details.
|
||||
|
|
@ -278,7 +278,7 @@ The generated package is a pip-installable Python package that registers as a `h
|
|||
|
||||
### Available Tools
|
||||
|
||||
`analyze`, `ask`, `get_document`, `list_documents`, `research`, `search`
|
||||
`cite`, `execute_code`, `get_document`, `list_documents`, `search`
|
||||
|
||||
### Example
|
||||
|
||||
|
|
@ -287,7 +287,7 @@ The generated package is a pip-installable Python package that registers as a `h
|
|||
haiku-rag create-skill \
|
||||
--name medic \
|
||||
--db /path/to/medic.lancedb \
|
||||
--tools search,ask \
|
||||
--tools search,cite \
|
||||
--config-file /path/to/haiku.rag.yaml \
|
||||
--description "Military medic knowledge base" \
|
||||
--preamble "You are a military medic expert."
|
||||
|
|
|
|||
|
|
@ -59,7 +59,6 @@ haiku-rag chat # Interactive conversation mode
|
|||
|
||||
- [Getting started](tutorial.md) - Tutorial
|
||||
- [Installation](installation.md) - Install haiku.rag with different providers
|
||||
- [Architecture](architecture.md) - System overview and data flow
|
||||
- [Configuration](configuration/index.md) - Environment variables and settings
|
||||
- [CLI](cli.md) - Command line interface usage
|
||||
- [Python](python.md) - Python API reference
|
||||
|
|
|
|||
|
|
@ -19,13 +19,10 @@ skill = create_skill(db_path=db_path, config=config)
|
|||
|
||||
| Tool | Purpose |
|
||||
|------|---------|
|
||||
| `analyze(question, document?, filter?)` | Answer analytical questions using code execution |
|
||||
|
||||
**Parameters:**
|
||||
|
||||
- `question` — The analytical question to answer.
|
||||
- `document` — Optional document ID or title to pre-load for analysis.
|
||||
- `filter` — Optional SQL WHERE clause to filter documents.
|
||||
| `search(query, limit?)` | Hybrid search (vector + full-text) with context expansion |
|
||||
| `list_documents()` | List all documents in the knowledge base |
|
||||
| `execute_code(code)` | Execute Python code in a sandboxed interpreter with VFS access |
|
||||
| `cite(chunk_ids)` | Register chunk IDs as citations for the current answer |
|
||||
|
||||
## State
|
||||
|
||||
|
|
@ -34,16 +31,16 @@ The skill manages an `AnalysisState` under the `"analysis"` namespace:
|
|||
```python
|
||||
class AnalysisState(BaseModel):
|
||||
document_filter: str | None = None
|
||||
analyses: list[AnalysisEntry] = []
|
||||
|
||||
class AnalysisEntry(BaseModel):
|
||||
question: str
|
||||
answer: str
|
||||
program: str | None = None
|
||||
executions: list[CodeExecutionEntry] = []
|
||||
citation_index: dict[str, Citation] = {}
|
||||
citations: list[list[str]] = []
|
||||
searches: dict[str, list[SearchResult]] = {}
|
||||
```
|
||||
|
||||
- **document_filter** — SQL WHERE clause applied to `analyze` calls (combined with any explicit `filter` parameter). Set this to scope analysis to specific documents.
|
||||
- **analyses** — Each `analyze` call appends an `AnalysisEntry` with the question, answer, and executed program.
|
||||
- **document_filter** — SQL WHERE clause applied to `search` and `list_documents` calls.
|
||||
- **executions** — Each `execute_code` call appends a `CodeExecutionEntry` with code, stdout, stderr, and success status.
|
||||
- **citation_index** / **citations** — Same per-turn citation tracking as the RAG skill.
|
||||
- **searches** — Search results from both the `search` tool and sandbox-internal searches.
|
||||
|
||||
## Usage with RAG Skill
|
||||
|
||||
|
|
@ -67,4 +64,4 @@ agent = Agent(
|
|||
)
|
||||
```
|
||||
|
||||
See the [Analysis Agent](../agents/analysis.md) documentation for details on how the underlying agent works.
|
||||
See the [Analysis Agent](../agents/analysis.md) documentation for details on how the underlying sandbox works.
|
||||
|
|
|
|||
|
|
@ -47,7 +47,7 @@ Use `create-skill` to generate a standalone skill package with an embedded datab
|
|||
haiku-rag create-skill \
|
||||
--name recipes \
|
||||
--db /path/to/recipes.lancedb \
|
||||
--tools search,ask \
|
||||
--tools search,cite \
|
||||
--description "Recipe knowledge base" \
|
||||
--preamble "You are a recipe expert."
|
||||
```
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
# RAG Skill
|
||||
|
||||
The RAG skill is the primary way to use haiku.rag tools. It bundles search, Q&A, document browsing, and research into a single skill with managed state.
|
||||
The RAG skill is the primary way to use haiku.rag tools. It bundles search, document browsing, and citation management into a single skill with managed state.
|
||||
|
||||
## `create_skill(db_path?, config?)`
|
||||
|
||||
|
|
@ -20,10 +20,9 @@ skill = create_skill(db_path=db_path, config=config)
|
|||
| Tool | Purpose |
|
||||
|------|---------|
|
||||
| `search(query, limit?)` | Hybrid search (vector + full-text) with context expansion |
|
||||
| `list_documents(limit?, offset?, filter?)` | Paginated document listing |
|
||||
| `list_documents()` | List all documents in the knowledge base |
|
||||
| `get_document(query)` | Retrieve a document by ID, title, or URI |
|
||||
| `ask(question)` | Q&A with citations via the QA agent |
|
||||
| `research(question)` | Deep multi-agent research producing comprehensive reports |
|
||||
| `cite(chunk_ids)` | Register chunk IDs as citations for the current answer |
|
||||
|
||||
## State
|
||||
|
||||
|
|
@ -31,17 +30,13 @@ The skill manages a `RAGState` under the `"rag"` namespace:
|
|||
|
||||
```python
|
||||
class RAGState(BaseModel):
|
||||
citations: list[Citation] = []
|
||||
qa_history: list[QAHistoryEntry] = []
|
||||
citation_index: dict[str, Citation] = {}
|
||||
citations: list[list[str]] = []
|
||||
document_filter: str | None = None
|
||||
searches: dict[str, list[SearchResult]] = {}
|
||||
documents: list[DocumentInfo] = []
|
||||
reports: list[ResearchEntry] = []
|
||||
```
|
||||
|
||||
- **citations** — Accumulated citations from `ask` calls, with sequential indexing across calls.
|
||||
- **qa_history** — Questions and answers from `ask` calls. Prior Q&A is used as context for follow-up questions when embeddings are similar.
|
||||
- **document_filter** — SQL WHERE clause applied to `search`, `list_documents`, `ask`, and `research` calls. Set this to scope queries to specific documents.
|
||||
- **citation_index** — All citations indexed by chunk ID (deduplicated across turns).
|
||||
- **citations** — Per-turn lists of chunk IDs registered via the `cite` tool.
|
||||
- **document_filter** — SQL WHERE clause applied to `search` and `list_documents` calls. Set this to scope queries to specific documents.
|
||||
- **searches** — Search results keyed by query string.
|
||||
- **documents** — Documents seen via `list_documents` or `get_document` (deduplicated by ID).
|
||||
- **reports** — Research reports from `research` calls.
|
||||
|
|
|
|||
|
|
@ -59,7 +59,6 @@ nav:
|
|||
- index.md
|
||||
- Getting started: tutorial.md
|
||||
- Installation: installation.md
|
||||
- Architecture: architecture.md
|
||||
- Configuration:
|
||||
- configuration/index.md
|
||||
- Providers: configuration/providers.md
|
||||
|
|
|
|||
Loading…
Reference in a new issue