255 lines
7.3 KiB
Markdown
255 lines
7.3 KiB
Markdown
# Command Line Interface
|
||
|
||
The `haiku-rag` CLI provides complete document management functionality.
|
||
|
||
!!! note
|
||
Global options (must be specified before the command):
|
||
|
||
- `--config` - Specify custom configuration file
|
||
- `--version` / `-v` - Show version and exit
|
||
|
||
Per-command options:
|
||
|
||
- `--db` - Specify custom database path
|
||
- `-h` - Show help for specific command
|
||
|
||
Example:
|
||
```bash
|
||
haiku-rag --config /path/to/config.yaml list
|
||
haiku-rag --config /path/to/config.yaml list --db /path/to/custom.db
|
||
haiku-rag add -h
|
||
```
|
||
|
||
## Document Management
|
||
|
||
### List Documents
|
||
|
||
```bash
|
||
haiku-rag list
|
||
```
|
||
|
||
Filter documents by properties:
|
||
```bash
|
||
# Filter by URI pattern
|
||
haiku-rag list --filter "uri LIKE '%arxiv%'"
|
||
|
||
# Filter by exact title
|
||
haiku-rag list --filter "title = 'My Document'"
|
||
|
||
# Combine multiple conditions
|
||
haiku-rag list --filter "uri LIKE '%.pdf' AND title LIKE '%paper%'"
|
||
```
|
||
|
||
### Add Documents
|
||
|
||
From text:
|
||
```bash
|
||
haiku-rag add "Your document content here"
|
||
|
||
# Attach metadata (repeat --meta for multiple entries)
|
||
haiku-rag add "Your document content here" --meta author=alice --meta topic=notes
|
||
```
|
||
|
||
From file or URL:
|
||
```bash
|
||
haiku-rag add-src /path/to/document.pdf
|
||
haiku-rag add-src https://example.com/article.html
|
||
|
||
# Optionally set a human‑readable title stored in the DB schema
|
||
haiku-rag add-src /mnt/data/doc1.pdf --title "Q3 Financial Report"
|
||
|
||
# Optionally attach metadata (repeat --meta). Values use JSON parsing if possible:
|
||
# numbers, booleans, null, arrays/objects; otherwise kept as strings.
|
||
haiku-rag add-src /mnt/data/doc1.pdf --meta source=manual --meta page_count=12 --meta published=true
|
||
```
|
||
|
||
From directory (recursively adds all supported files):
|
||
```bash
|
||
haiku-rag add-src /path/to/documents/
|
||
```
|
||
|
||
!!! note
|
||
When adding a directory, the same content filters configured for [file monitoring](configuration.md#filtering-monitored-files) are applied. This means `ignore_patterns` and `include_patterns` from your configuration will be used to filter which files are added.
|
||
|
||
!!! note
|
||
As you add documents to `haiku.rag` the database keeps growing. By default, LanceDB supports versioning
|
||
of your data. Create/update operations are atomic‑feeling: if anything fails during chunking or embedding,
|
||
the database rolls back to the pre‑operation snapshot using LanceDB table versioning. You can optimize and
|
||
compact the database by running the [vacuum](#vacuum-optimize-and-cleanup) command.
|
||
|
||
### Get Document
|
||
|
||
```bash
|
||
haiku-rag get 3f4a... # document ID
|
||
```
|
||
|
||
### Delete Document
|
||
|
||
```bash
|
||
haiku-rag delete 3f4a... # document ID
|
||
haiku-rag rm 3f4a... # alias
|
||
```
|
||
|
||
Use this when you want to change things like the embedding model or chunk size for example.
|
||
|
||
## Search
|
||
|
||
Basic search:
|
||
```bash
|
||
haiku-rag search "machine learning"
|
||
```
|
||
|
||
With options:
|
||
```bash
|
||
haiku-rag search "python programming" --limit 10
|
||
```
|
||
|
||
With filters (filter by document properties):
|
||
```bash
|
||
# Filter by URI pattern
|
||
haiku-rag search "neural networks" --filter "uri LIKE '%arxiv%'"
|
||
|
||
# Filter by exact title
|
||
haiku-rag search "transformers" --filter "title = 'Deep Learning Guide'"
|
||
|
||
# Combine multiple conditions
|
||
haiku-rag search "AI" --filter "uri LIKE '%.pdf' AND title LIKE '%paper%'"
|
||
```
|
||
|
||
## Question Answering
|
||
|
||
Ask questions about your documents:
|
||
```bash
|
||
haiku-rag ask "Who is the author of haiku.rag?"
|
||
```
|
||
|
||
Ask questions with citations showing source documents:
|
||
```bash
|
||
haiku-rag ask "Who is the author of haiku.rag?" --cite
|
||
```
|
||
|
||
Use deep QA for complex questions (multi-agent decomposition):
|
||
```bash
|
||
haiku-rag ask "What are the main features and architecture of haiku.rag?" --deep --cite
|
||
```
|
||
|
||
Show verbose output with deep QA:
|
||
```bash
|
||
haiku-rag ask "What are the main features and architecture of haiku.rag?" --deep --verbose
|
||
```
|
||
|
||
The QA agent will search your documents for relevant information and provide a comprehensive answer. With `--cite`, responses include citations showing which documents were used. With `--deep`, the question is decomposed into sub-questions that are answered in parallel before synthesizing a final answer. With `--verbose` (only with `--deep`), you'll see the planning, searching, evaluation, and synthesis steps as they happen.
|
||
When available, citations use the document title; otherwise they fall back to the URI.
|
||
|
||
## Research
|
||
|
||
Run the multi-step research graph:
|
||
|
||
```bash
|
||
haiku-rag research "How does haiku.rag organize and query documents?"
|
||
```
|
||
|
||
With verbose output to see progress:
|
||
|
||
```bash
|
||
haiku-rag research "How does haiku.rag organize and query documents?" --verbose
|
||
```
|
||
|
||
Flags:
|
||
- `--verbose`: Show planning, searching previews, evaluation summary, and stop reason
|
||
|
||
Research parameters like `max_iterations`, `confidence_threshold`, and `max_concurrency` are configured in your [configuration file](configuration.md) under the `research` section.
|
||
|
||
When `--verbose` is set, the CLI consumes the research graph's AG-UI event stream, displaying step events and activity snapshots as agents progress through planning, search, evaluation, and synthesis. Without `--verbose`, only the final research report is displayed.
|
||
|
||
If you build your own integration, import `stream_graph` from `haiku.rag.graph.agui` to access AG-UI events (`STEP_STARTED`, `ACTIVITY_SNAPSHOT`, `STATE_SNAPSHOT`, `RUN_FINISHED`, etc.) and render them however you like while the graph is running.
|
||
|
||
## Server
|
||
|
||
Start services (requires at least one flag):
|
||
```bash
|
||
# MCP server only (HTTP transport)
|
||
haiku-rag serve --mcp
|
||
|
||
# MCP server (stdio transport)
|
||
haiku-rag serve --mcp --stdio
|
||
|
||
# File monitoring only
|
||
haiku-rag serve --monitor
|
||
|
||
# AG-UI server only
|
||
haiku-rag serve --agui
|
||
|
||
# Multiple services
|
||
haiku-rag serve --monitor --mcp
|
||
haiku-rag serve --monitor --agui
|
||
haiku-rag serve --mcp --agui
|
||
|
||
# All services
|
||
haiku-rag serve --monitor --mcp --agui
|
||
|
||
# Custom MCP port
|
||
haiku-rag serve --mcp --mcp-port 9000
|
||
```
|
||
|
||
See [Server Mode](server.md) for details on available services.
|
||
|
||
## Settings
|
||
|
||
View current configuration settings:
|
||
```bash
|
||
haiku-rag settings
|
||
```
|
||
|
||
## Maintenance
|
||
|
||
### Info (Read-only)
|
||
|
||
Display database metadata without upgrading or modifying it:
|
||
|
||
```bash
|
||
haiku-rag info [--db /path/to/your.lancedb]
|
||
```
|
||
|
||
Shows:
|
||
- path to the database
|
||
- stored haiku.rag version (from settings)
|
||
- embeddings provider/model and vector dimension
|
||
- number of documents
|
||
- table versions per table (documents, chunks)
|
||
|
||
At the end, a separate “Versions” section lists runtime package versions:
|
||
- haiku.rag
|
||
- lancedb
|
||
- docling
|
||
|
||
### Vacuum (Optimize and Cleanup)
|
||
|
||
Reduce disk usage by optimizing and pruning old table versions across all tables:
|
||
|
||
```bash
|
||
haiku-rag vacuum
|
||
```
|
||
|
||
**Automatic Cleanup:** Vacuum runs automatically in the background after document operations. By default, it removes versions older than 1 day (configurable via `storage.vacuum_retention_seconds`), preserving recent versions for concurrent connections. Manual vacuum can be useful for cleanup after bulk operations or to free disk space immediately.
|
||
|
||
### Rebuild Database
|
||
|
||
Rebuild the database by deleting all chunks & embeddings and re-indexing all documents. This is useful
|
||
when want to switch embeddings provider or model:
|
||
|
||
```bash
|
||
haiku-rag rebuild
|
||
```
|
||
|
||
### Download Models
|
||
|
||
Download required runtime models:
|
||
|
||
```bash
|
||
haiku-rag download-models
|
||
```
|
||
|
||
This command:
|
||
- Downloads Docling OCR/conversion models (no-op if already present).
|
||
- Pulls Ollama models referenced in your configuration (embeddings, QA, research, rerank).
|