Update docs

This commit is contained in:
Yiorgis Gozadinos 2025-12-15 13:43:24 +02:00
parent 557a296341
commit 6ee0d74f92
No known key found for this signature in database
2 changed files with 80 additions and 190 deletions

View file

@ -1,156 +1,57 @@
# Agents # Agents
Three agentic flows are provided by haiku.rag: Two agentic flows are provided by haiku.rag:
- Simple QA Agent — a focused question answering agent - **Simple QA Agent** — a focused question answering agent
- Deep QA Agent — multi-agent question decomposition for complex questions - **Research Graph** — a multi-step research workflow with question decomposition
- Research MultiAgent — a multistep, analyzable research workflow
For an interactive example using Pydantic AI and AG-UI, see the [Interactive Research Assistant](https://github.com/ggozad/haiku.rag/tree/main/examples/ag-ui-research) example ([demo video](https://vimeo.com/1128874386)). The demo uses a knowledge base containing haiku.rag's code and documentation. For an interactive example using Pydantic AI and AG-UI, see the [Interactive Research Assistant](https://github.com/ggozad/haiku.rag/tree/main/examples/ag-ui-research) example ([demo video](https://vimeo.com/1128874386)).
See [QA and Research Configuration](configuration/qa-research.md) for configuring model, iterations, concurrency, and other settings. See [QA and Research Configuration](configuration/qa-research.md) for configuring model, iterations, concurrency, and other settings.
## Simple QA Agent
### Simple QA Agent
The simple QA agent answers a single question using the knowledge base. It retrieves relevant chunks, optionally expands context around them, and asks the model to answer strictly based on that context. The simple QA agent answers a single question using the knowledge base. It retrieves relevant chunks, optionally expands context around them, and asks the model to answer strictly based on that context.
Key points: Key points:
- Uses a single `search_documents` tool to fetch relevant chunks - Uses a single `search_documents` tool to fetch relevant chunks
- Can be run with or without inline citations in the prompt (citations prefer - Can be run with or without inline citations in the prompt
document titles when present, otherwise URIs)
- Returns a plain string answer - Returns a plain string answer
Python usage: **CLI usage:**
```bash
haiku-rag ask "What is climate change?"
# With citations
haiku-rag ask "What is climate change?" --cite
# Deep mode (uses research graph with optimized settings)
haiku-rag ask "What are the main features of haiku.rag?" --deep
```
**Python usage:**
```python ```python
from haiku.rag.client import HaikuRAG from haiku.rag.client import HaikuRAG
from haiku.rag.qa.agent import QuestionAnswerAgent from haiku.rag.qa.agent import QuestionAnswerAgent
async with HaikuRAG(path_to_db) as client: async with HaikuRAG(path_to_db) as client:
# Choose a provider and model (see Configuration for env defaults)
agent = QuestionAnswerAgent( agent = QuestionAnswerAgent(
client=client, client=client,
provider="openai", # or "ollama", "vllm", etc. provider="openai",
model="gpt-4o-mini", model="gpt-4o-mini",
use_citations=False, # set True to bias prompt towards citing sources use_citations=False,
) )
answer = await agent.answer("What is climate change?") answer = await agent.answer("What is climate change?")
print(answer) print(answer)
``` ```
### Deep QA Agent ## Research Graph
Deep QA is a multi-agent system that decomposes complex questions into sub-questions, answers them in batches, evaluates sufficiency, and iterates if needed before synthesizing a final answer. It's lighter than the full research workflow but more powerful than the simple QA agent. The research workflow is implemented as a typed pydantic-graph. It plans, searches (in parallel batches), evaluates, and synthesizes into a final report.
```mermaid
---
title: Deep QA graph
---
stateDiagram-v2
[*] --> plan
plan --> get_batch
get_batch --> search_one: Has questions (map)
get_batch --> synthesize: No questions
search_one --> collect_answers
collect_answers --> decide
decide --> get_batch: Continue QA
decide --> synthesize: Done with QA
synthesize --> [*]
```
Key nodes:
- **plan**: Decomposes the question into focused sub-questions using a presearch tool
- **get_batch**: Retrieves remaining sub-questions for the current iteration
- **search_one**: Answers a single sub-question using the knowledge base (mapped in parallel)
- **collect_answers**: Aggregates search results from parallel executions
- **decide**: Evaluates if sufficient information has been gathered or if more iterations are needed
- **synthesize**: Generates the final comprehensive answer from all gathered information
Key differences from Research:
- **Simpler evaluation**: Uses sufficiency check (not confidence + insight analysis)
- **Direct answers**: Returns just the answer (not a full research report)
- **Question-focused**: Optimized for answering specific questions, not open-ended research
- **Supports citations**: Can include inline source citations like `[document.md]`
- **Configurable iterations**: Control max_iterations (default: 2) and max_concurrency (default: 1)
Note on parallel execution:
- The `search_one` node is mapped over all questions in a batch
- Parallelism is controlled via `max_concurrency`
- All questions in an iteration are processed before evaluation
CLI usage:
```bash
# Deep QA without citations
haiku-rag ask "What are the main features of haiku.rag?" --deep
# Deep QA with citations
haiku-rag ask "What are the main features of haiku.rag?" --deep --cite
```
Python usage:
```python
from haiku.rag.client import HaikuRAG
from haiku.rag.config import Config
from haiku.rag.graph.deep_qa.dependencies import DeepQAContext
from haiku.rag.graph.deep_qa.graph import build_deep_qa_graph
from haiku.rag.graph.deep_qa.state import DeepQADeps, DeepQAState
async with HaikuRAG(path_to_db) as client:
# Use global config (recommended)
graph = build_deep_qa_graph(config=Config)
context = DeepQAContext(
original_question="What are the main features of haiku.rag?",
use_citations=True
)
state = DeepQAState.from_config(context=context, config=Config)
deps = DeepQADeps(client=client)
result = await graph.run(
state=state,
deps=deps
)
print(result.answer)
print(result.sources)
```
Alternative usage with custom config:
```python
# Create a custom config with different settings
from haiku.rag.config.models import AppConfig, QAConfig
custom_config = AppConfig(
qa=QAConfig(
provider="openai",
model="gpt-4o-mini",
max_sub_questions=5,
max_iterations=3,
max_concurrency=2,
)
)
graph = build_deep_qa_graph(config=custom_config)
context = DeepQAContext(
original_question="What are the main features of haiku.rag?",
use_citations=True
)
state = DeepQAState.from_config(context=context, config=custom_config)
deps = DeepQADeps(client=client)
result = await graph.run(state=state, deps=deps)
```
### Research Graph
The research workflow is implemented as a typed pydanticgraph. It plans, searches (in parallel batches), evaluates, and synthesizes into a final report — with clear stop conditions and shared state.
```mermaid ```mermaid
--- ---
@ -162,47 +63,49 @@ stateDiagram-v2
get_batch --> search_one: Has questions (map) get_batch --> search_one: Has questions (map)
get_batch --> synthesize: No questions get_batch --> synthesize: No questions
search_one --> collect_answers search_one --> collect_answers
collect_answers --> analyze_insights collect_answers --> decide
analyze_insights --> decide
decide --> get_batch: Continue research decide --> get_batch: Continue research
decide --> synthesize: Done researching decide --> synthesize: Done researching
synthesize --> [*] synthesize --> [*]
``` ```
Key nodes: **Key nodes:**
- **plan**: Builds up to 3 standalone subquestions (uses an internal presearch tool) - **plan**: Builds up to 3 standalone sub-questions (uses an internal presearch tool)
- **get_batch**: Retrieves remaining subquestions for the current iteration - **get_batch**: Retrieves remaining sub-questions for the current iteration
- **search_one**: Answers a single subquestion using the KB with minimal, verbatim context (mapped in parallel) - **search_one**: Answers a single sub-question using the KB (mapped in parallel)
- **collect_answers**: Aggregates search results from parallel executions - **collect_answers**: Aggregates search results from parallel executions
- **analyze_insights**: Synthesizes fresh insights, updates gaps, and suggests new sub-questions - **decide**: Evaluates confidence and determines whether to continue or synthesize
- **decide**: Checks sufficiency/confidence thresholds and determines whether to continue research
- **synthesize**: Generates a final structured research report - **synthesize**: Generates a final structured research report
Primary models: **Primary models:**
- `SearchAnswer` — one per subquestion (query, answer, context, sources) - `SearchAnswer` — one per sub-question (query, answer, confidence, citations)
- `InsightRecord` / `GapRecord` — structured tracking of findings and open issues - `EvaluationResult` — confidence score, new questions, sufficiency assessment
- `InsightAnalysis` — output of the analysis stage (insights, gaps, commentary)
- `EvaluationResult` — insights, new questions, sufficiency, confidence
- `ResearchReport` — final report (title, executive summary, findings, conclusions, …) - `ResearchReport` — final report (title, executive summary, findings, conclusions, …)
Note on parallel execution: **Parallel execution:**
- The `search_one` node is mapped over all questions in a batch - The `search_one` node is mapped over all questions in a batch
- Parallelism is controlled via `max_concurrency` - Parallelism is controlled via `max_concurrency`
- Analysis and decision nodes process results after each batch completes - Decision nodes process results after each batch completes
CLI usage: ### CLI Usage
```bash ```bash
# Basic usage (uses config from file or defaults) # Basic usage
haiku-rag research "How does haiku.rag organize and query documents?"
# With verbose output (shows progress)
haiku-rag research "How does haiku.rag organize and query documents?" --verbose haiku-rag research "How does haiku.rag organize and query documents?" --verbose
# With custom config file # With document filter
haiku-rag --config my-research-config.yaml research "How does haiku.rag organize and query documents?" --verbose haiku-rag research "What are the key findings?" --filter "uri LIKE '%report%'"
``` ```
Python usage (blocking result): ### Python Usage
**Basic example:**
```python ```python
from haiku.rag.client import HaikuRAG from haiku.rag.client import HaikuRAG
@ -212,41 +115,25 @@ from haiku.rag.graph.research.graph import build_research_graph
from haiku.rag.graph.research.state import ResearchDeps, ResearchState from haiku.rag.graph.research.state import ResearchDeps, ResearchState
async with HaikuRAG(path_to_db) as client: async with HaikuRAG(path_to_db) as client:
# Use global config (recommended)
graph = build_research_graph(config=Config) graph = build_research_graph(config=Config)
question = "What are the main drivers and trends of global temperature anomalies since 1990?" context = ResearchContext(original_question="What are the main features?")
context = ResearchContext(original_question=question)
state = ResearchState.from_config(context=context, config=Config) state = ResearchState.from_config(context=context, config=Config)
deps = ResearchDeps(client=client) deps = ResearchDeps(client=client)
result = await graph.run( report = await graph.run(state=state, deps=deps)
state=state,
deps=deps,
)
report = result
print(report.title) print(report.title)
print(report.executive_summary) print(report.executive_summary)
``` ```
### Filtering Documents **With custom config:**
Both Research and Deep QA graphs support restricting searches to specific documents via the `search_filter` parameter. Set it to a SQL WHERE clause before running:
```python
state = ResearchState.from_config(context=context, config=Config)
# Only search documents with these IDs
state.search_filter = "id IN ('doc-123', 'doc-456')"
result = await graph.run(state=state, deps=deps)
```
The filter applies to all search operations in the graph (context gathering and sub-question searches). See [Filtering Search Results](python.md#filtering-search-results) for available filter columns and syntax.
Alternative usage with custom config:
```python ```python
from haiku.rag.client import HaikuRAG
from haiku.rag.config.models import AppConfig, ResearchConfig from haiku.rag.config.models import AppConfig, ResearchConfig
from haiku.rag.graph.research.dependencies import ResearchContext
from haiku.rag.graph.research.graph import build_research_graph
from haiku.rag.graph.research.state import ResearchDeps, ResearchState
custom_config = AppConfig( custom_config = AppConfig(
research=ResearchConfig( research=ResearchConfig(
@ -258,28 +145,28 @@ custom_config = AppConfig(
) )
) )
graph = build_research_graph(config=custom_config) async with HaikuRAG(path_to_db) as client:
context = ResearchContext(original_question=question) graph = build_research_graph(config=custom_config)
state = ResearchState.from_config(context=context, config=custom_config) context = ResearchContext(original_question="What are the main features?")
deps = ResearchDeps(client=client) state = ResearchState.from_config(context=context, config=custom_config)
deps = ResearchDeps(client=client)
result = await graph.run(state=state, deps=deps) report = await graph.run(state=state, deps=deps)
``` ```
Python usage (streamed AG-UI events): **Streaming AG-UI events:**
```python ```python
from haiku.rag.graph.agui import stream_graph
from haiku.rag.client import HaikuRAG from haiku.rag.client import HaikuRAG
from haiku.rag.config import Config from haiku.rag.config import Config
from haiku.rag.graph.agui import stream_graph
from haiku.rag.graph.research.dependencies import ResearchContext from haiku.rag.graph.research.dependencies import ResearchContext
from haiku.rag.graph.research.graph import build_research_graph from haiku.rag.graph.research.graph import build_research_graph
from haiku.rag.graph.research.state import ResearchDeps, ResearchState from haiku.rag.graph.research.state import ResearchDeps, ResearchState
async with HaikuRAG(path_to_db) as client: async with HaikuRAG(path_to_db) as client:
graph = build_research_graph(config=Config) graph = build_research_graph(config=Config)
question = "What are the main drivers and trends of global temperature anomalies since 1990?" context = ResearchContext(original_question="What are the main features?")
context = ResearchContext(original_question=question)
state = ResearchState.from_config(context=context, config=Config) state = ResearchState.from_config(context=context, config=Config)
deps = ResearchDeps(client=client) deps = ResearchDeps(client=client)
@ -287,21 +174,25 @@ async with HaikuRAG(path_to_db) as client:
if event["type"] == "STEP_STARTED": if event["type"] == "STEP_STARTED":
print(f"Starting step: {event['stepName']}") print(f"Starting step: {event['stepName']}")
elif event["type"] == "ACTIVITY_SNAPSHOT": elif event["type"] == "ACTIVITY_SNAPSHOT":
# Activity events include structured data alongside messages content = event["content"]
content = event['content']
print(f" {content['message']}") print(f" {content['message']}")
if "confidence" in content:
# Different activity types have different structured fields
if 'confidence' in content:
print(f" Confidence: {content['confidence']:.0%}") print(f" Confidence: {content['confidence']:.0%}")
if 'sub_questions' in content:
for q in content['sub_questions']:
print(f" - {q}")
if 'insights' in content:
print(f" New insights: {len(content['insights'])}")
elif event["type"] == "RUN_FINISHED": elif event["type"] == "RUN_FINISHED":
print("\nResearch complete!\n") report = event["result"]
result = event["result"] print(report["executive_summary"])
print(result["title"])
print(result["executive_summary"])
``` ```
### Filtering Documents
Restrict searches to specific documents via the `search_filter` parameter:
```python
# Set filter before running the graph
state = ResearchState.from_config(context=context, config=Config)
state.search_filter = "id IN ('doc-123', 'doc-456')"
report = await graph.run(state=state, deps=deps)
```
The filter applies to all search operations in the graph. See [Filtering Search Results](python.md#filtering-search-results) for available filter columns and syntax.

View file

@ -239,8 +239,7 @@ data: {"type":"RUN_FINISHED","threadId":"abc123","runId":"xyz789","result":{"tit
- Additional structured fields depending on the activity type: - Additional structured fields depending on the activity type:
- **Planning**: `sub_questions` (list of strings) - **Planning**: `sub_questions` (list of strings)
- **Searching**: `query` (string), `confidence` (float, on completion), `error` (string, on failure) - **Searching**: `query` (string), `confidence` (float, on completion), `error` (string, on failure)
- **Analyzing**: `insights` (list of insight objects), `gaps` (list of gap objects), `resolved_gaps` (list of strings) - **Evaluating**: `confidence` (float), `is_sufficient` (boolean), `new_questions` (list of strings)
- **Evaluating**: `confidence` (float), `is_sufficient` (boolean) for research; `is_sufficient` (boolean), `iterations` (int) for deep QA
The `message` field is always present for simple rendering, while structured fields enable richer UI features like displaying lists, charts, and detailed status information. The `message` field is always present for simple rendering, while structured fields enable richer UI features like displaying lists, charts, and detailed status information.