haiku.rag/evaluations/README.md

13 lines
374 B
Markdown

# Haiku RAG - Evaluations
Internal benchmarking and evaluation scripts for haiku.rag.
This package is not published to PyPI and is only used for development and testing purposes.
## Overview
Contains evaluation scripts for benchmarking RAG performance using datasets like:
- RepliQA
- WiX
Uses `pydantic-evals` for evaluation framework and `logfire` for observability.