haiku.rag/docs/agents.md
Yiorgis Gozadinos fa93226a28
Update docs
2025-11-06 11:02:33 +02:00

8.5 KiB
Raw Blame History

Agents

Three agentic flows are provided by haiku.rag:

  • Simple QA Agent — a focused question answering agent
  • Deep QA Agent — multi-agent question decomposition for complex questions
  • Research MultiAgent — a multistep, analyzable research workflow

For an interactive example using Pydantic AI and AG-UI, see the Interactive Research Assistant example (demo video). The demo uses a knowledge base containing haiku.rag's code and documentation.

Simple QA Agent

The simple QA agent answers a single question using the knowledge base. It retrieves relevant chunks, optionally expands context around them, and asks the model to answer strictly based on that context.

Key points:

  • Uses a single search_documents tool to fetch relevant chunks
  • Can be run with or without inline citations in the prompt (citations prefer document titles when present, otherwise URIs)
  • Returns a plain string answer

Python usage:

from haiku.rag.client import HaikuRAG
from haiku.rag.qa.agent import QuestionAnswerAgent

async with HaikuRAG(path_to_db) as client:
    # Choose a provider and model (see Configuration for env defaults)
    agent = QuestionAnswerAgent(
        client=client,
        provider="openai",  # or "ollama", "vllm", etc.
        model="gpt-4o-mini",
        use_citations=False,  # set True to bias prompt towards citing sources
    )

    answer = await agent.answer("What is climate change?")
    print(answer)

Deep QA Agent

Deep QA is a multi-agent system that decomposes complex questions into sub-questions, answers them in batches, evaluates sufficiency, and iterates if needed before synthesizing a final answer. It's lighter than the full research workflow but more powerful than the simple QA agent.

---
title: Deep QA graph
---
stateDiagram-v2
  [*] --> plan
  plan --> get_batch
  get_batch --> search_one: Has questions (map)
  get_batch --> synthesize: No questions
  search_one --> collect_answers
  collect_answers --> decide
  decide --> get_batch: Continue QA
  decide --> synthesize: Done with QA
  synthesize --> [*]

Key nodes:

  • plan: Decomposes the question into focused sub-questions using a presearch tool
  • get_batch: Retrieves remaining sub-questions for the current iteration
  • search_one: Answers a single sub-question using the knowledge base (mapped in parallel)
  • collect_answers: Aggregates search results from parallel executions
  • decide: Evaluates if sufficient information has been gathered or if more iterations are needed
  • synthesize: Generates the final comprehensive answer from all gathered information

Key differences from Research:

  • Simpler evaluation: Uses sufficiency check (not confidence + insight analysis)
  • Direct answers: Returns just the answer (not a full research report)
  • Question-focused: Optimized for answering specific questions, not open-ended research
  • Supports citations: Can include inline source citations like [document.md]
  • Configurable iterations: Control max_iterations (default: 2) and max_concurrency (default: 1)

Note on parallel execution:

  • The search_one node is mapped over all questions in a batch
  • Parallelism is controlled via max_concurrency using asyncio.Semaphore
  • All questions in an iteration are processed before evaluation

CLI usage:

# Deep QA without citations
haiku-rag ask "What are the main features of haiku.rag?" --deep

# Deep QA with citations
haiku-rag ask "What are the main features of haiku.rag?" --deep --cite

Python usage:

from haiku.rag.client import HaikuRAG
from haiku.rag.qa.deep.dependencies import DeepQAContext
from haiku.rag.qa.deep.graph import build_deep_qa_graph
from haiku.rag.qa.deep.state import DeepQADeps, DeepQAState

async with HaikuRAG(path_to_db) as client:
    graph = build_deep_qa_graph(
        provider="openai",
        model="gpt-4o-mini"
    )
    context = DeepQAContext(
        original_question="What are the main features of haiku.rag?",
        use_citations=True
    )
    state = DeepQAState(
        context=context,
        max_sub_questions=3,
        max_iterations=2,
        max_concurrency=1
    )
    deps = DeepQADeps(client=client)

    result = await graph.run(
        state=state,
        deps=deps
    )

    print(result.answer)
    print(result.sources)

Research Graph

The research workflow is implemented as a typed pydanticgraph. It plans, searches (in parallel batches), evaluates, and synthesizes into a final report — with clear stop conditions and shared state.

---
title: Research graph
---
stateDiagram-v2
  [*] --> plan
  plan --> get_batch
  get_batch --> search_one: Has questions (map)
  get_batch --> synthesize: No questions
  search_one --> collect_answers
  collect_answers --> analyze_insights
  analyze_insights --> decide
  decide --> get_batch: Continue research
  decide --> synthesize: Done researching
  synthesize --> [*]

Key nodes:

  • plan: Builds up to 3 standalone subquestions (uses an internal presearch tool)
  • get_batch: Retrieves remaining subquestions for the current iteration
  • search_one: Answers a single subquestion using the KB with minimal, verbatim context (mapped in parallel)
  • collect_answers: Aggregates search results from parallel executions
  • analyze_insights: Synthesizes fresh insights, updates gaps, and suggests new sub-questions
  • decide: Checks sufficiency/confidence thresholds and determines whether to continue research
  • synthesize: Generates a final structured research report

Primary models:

  • SearchAnswer — one per subquestion (query, answer, context, sources)
  • InsightRecord / GapRecord — structured tracking of findings and open issues
  • InsightAnalysis — output of the analysis stage (insights, gaps, commentary)
  • EvaluationResult — insights, new questions, sufficiency, confidence
  • ResearchReport — final report (title, executive summary, findings, conclusions, …)

Note on parallel execution:

  • The search_one node is mapped over all questions in a batch
  • Parallelism is controlled via max_concurrency using asyncio.Semaphore
  • Analysis and decision nodes process results after each batch completes

CLI usage:

haiku-rag research "How does haiku.rag organize and query documents?" \
  --max-iterations 2 \
  --confidence-threshold 0.8 \
  --max-concurrency 3 \
  --verbose

Python usage (blocking result):

from haiku.rag.client import HaikuRAG
from haiku.rag.research.dependencies import ResearchContext
from haiku.rag.research.graph import build_research_graph
from haiku.rag.research.state import ResearchDeps, ResearchState

async with HaikuRAG(path_to_db) as client:
    graph = build_research_graph(
        provider="openai",
        model="gpt-4o-mini"
    )
    question = "What are the main drivers and trends of global temperature anomalies since 1990?"
    state = ResearchState(
        context=ResearchContext(original_question=question),
        max_iterations=2,
        confidence_threshold=0.8,
        max_concurrency=2,
    )
    deps = ResearchDeps(client=client)

    result = await graph.run(
        state=state,
        deps=deps,
    )

    report = result
    print(report.title)
    print(report.executive_summary)

Python usage (streamed events):

from haiku.rag.client import HaikuRAG
from haiku.rag.research.dependencies import ResearchContext
from haiku.rag.research.graph import build_research_graph
from haiku.rag.research.state import ResearchDeps, ResearchState
from haiku.rag.research.stream import stream_research_graph

async with HaikuRAG(path_to_db) as client:
    graph = build_research_graph(
        provider="openai",
        model="gpt-4o-mini"
    )
    question = "What are the main drivers and trends of global temperature anomalies since 1990?"
    state = ResearchState(
        context=ResearchContext(original_question=question),
        max_iterations=2,
        confidence_threshold=0.8,
        max_concurrency=2,
    )
    deps = ResearchDeps(client=client)

    async for event in stream_research_graph(
        graph,
        state,
        deps,
    ):
        if event.type == "log":
            iteration = event.state.iterations if event.state else state.iterations
            print(f"[{iteration}] {event.message}")
        elif event.type == "report":
            print("\nResearch complete!\n")
            print(event.report.title)
            print(event.report.executive_summary)