Update docs
This commit is contained in:
parent
7cc4e13b57
commit
87f28ed96b
2 changed files with 10 additions and 2 deletions
|
|
@ -3,7 +3,9 @@
|
||||||
We use the [repliqa](https://huggingface.co/datasets/ServiceNow/repliqa) dataset for the evaluation of `haiku.rag`.
|
We use the [repliqa](https://huggingface.co/datasets/ServiceNow/repliqa) dataset for the evaluation of `haiku.rag`.
|
||||||
|
|
||||||
You can perform your own evaluations using as example the script found at
|
You can perform your own evaluations using as example the script found at
|
||||||
`tests/generate_benchmark_db.py`.
|
`tests/generate_benchmark_db.py`. The evaluation flow is orchestrated with
|
||||||
|
[`pydantic-evals`](https://github.com/pydantic/pydantic-ai/tree/main/libs/pydantic-evals),
|
||||||
|
which we leverage for dataset management, scoring, and report generation.
|
||||||
|
|
||||||
## Recall
|
## Recall
|
||||||
|
|
||||||
|
|
@ -24,7 +26,9 @@ The recall obtained is ~0.79 for matching in the top result, raising to ~0.91 fo
|
||||||
|
|
||||||
## Question/Answer evaluation
|
## Question/Answer evaluation
|
||||||
|
|
||||||
Again using the same dataset, we use a QA agent to answer the question. In addition we use an LLM judge (using the Ollama `qwen3`) to evaluate whether the answer is correct or not. The obtained accuracy is as follows:
|
Again using the same dataset, we use a QA agent to answer the question.
|
||||||
|
`pydantic-evals` runs each case and coordinates an LLM judge (Ollama `qwen3`) to
|
||||||
|
determine whether the answer is correct. The obtained accuracy is as follows:
|
||||||
|
|
||||||
| Embedding Model | QA Model | Accuracy | Reranker |
|
| Embedding Model | QA Model | Accuracy | Reranker |
|
||||||
|------------------------------------|-----------------------------------|-----------|------------------------|
|
|------------------------------------|-----------------------------------|-----------|------------------------|
|
||||||
|
|
|
||||||
|
|
@ -205,6 +205,7 @@ async def run_qa_benchmark(k: int | None = None):
|
||||||
)
|
)
|
||||||
if equivalence.value:
|
if equivalence.value:
|
||||||
passing_cases += 1
|
passing_cases += 1
|
||||||
|
|
||||||
console.print("")
|
console.print("")
|
||||||
|
|
||||||
if report.failures:
|
if report.failures:
|
||||||
|
|
@ -215,6 +216,9 @@ async def run_qa_benchmark(k: int | None = None):
|
||||||
console.print(f"Error: {failure.error_message}")
|
console.print(f"Error: {failure.error_message}")
|
||||||
console.print("")
|
console.print("")
|
||||||
|
|
||||||
|
console.print(
|
||||||
|
f"[green]Accuracy: {(passing_cases / total_processed):.4f}[/green]"
|
||||||
|
)
|
||||||
total_cases = total_processed
|
total_cases = total_processed
|
||||||
accuracy = passing_cases / total_cases if total_cases > 0 else 0
|
accuracy = passing_cases / total_cases if total_cases > 0 else 0
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue