Update orb qa accuracy

This commit is contained in:
Yiorgis Gozadinos 2026-01-25 16:27:46 +02:00
parent c7a9ad8583
commit 50a1a6171c
No known key found for this signature in database

View file

@ -110,7 +110,6 @@ We benchmark both the plain text version (HTML stripped, no structure) and HTML
[HotpotQA](https://huggingface.co/datasets/hotpotqa/hotpot_qa) is a multi-hop question answering dataset requiring reasoning over multiple Wikipedia paragraphs. Each question requires evidence from 2+ documents, making it ideal for testing retrieval and reasoning capabilities. We use MAP for retrieval evaluation since queries have multiple relevant documents.
*Results from v0.20.2*
QA accuracy is evaluated over 2000 "hard" questions from the validation dataset.
### Retrieval (MAP)
@ -137,3 +136,9 @@ QA accuracy is evaluated over 2000 "hard" questions from the validation dataset.
| Embedding Model | MAP | VLM |
|----------------------|--------|----------------------|
| `qwen3-embedding:4b` | 0.9626 | Ollama / ministral-3 |
### QA Accuracy
| Embedding Model | QA Model | Accuracy | VLM |
|----------------------|-----------------------------|----------|----------------------|
| `qwen3-embedding:4b` | `gpt-oss:20b` - no thinking | 0.912 | Ollama / ministral-3 |