Update benchmarks with reranking

This commit is contained in:
Yiorgis Gozadinos 2025-07-19 19:02:47 +03:00
parent a6595b1597
commit c1d5334804
No known key found for this signature in database
3 changed files with 36 additions and 21 deletions

View file

@ -12,17 +12,19 @@ In order to calculate recall, we load the `News Stories` from `repliqa_3` which
The recall obtained is ~0.73 for matching in the top result, raising to ~0.75 for the top 3 results. The recall obtained is ~0.73 for matching in the top result, raising to ~0.75 for the top 3 results.
| Model | Document in top 1 | Document in top 3 | | Model | Document in top 1 | Document in top 3 | Reranker |
|---------------------------------------|-------------------|-------------------| |---------------------------------------|-------------------|-------------------|----------------------|
| Ollama / `mxbai-embed-large` | 0.77 | 0.89 | | Ollama / `mxbai-embed-large` | 0.77 | 0.89 | None |
| Ollama / `nomic-embed-text` | 0.74 | 0.88 | | Ollama / `mxbai-embed-large` | 0.81 | 0.91 | mxbai-rerank-base-v2 |
| OpenAI / `text-embeddings-3-small` | 0.75 | 0.88 | | Ollama / `nomic-embed-text` | 0.74 | 0.88 | None |
| OpenAI / `text-embeddings-3-small` | 0.75 | 0.88 | None |
## Question/Answer evaluation ## Question/Answer evaluation
Again using the same dataset, we use a QA agent to answer the question. In addition we use an LLM judge (using the Ollama `qwen3`) to evaluate whether the answer is correct or not. The obtained accuracy is as follows: Again using the same dataset, we use a QA agent to answer the question. In addition we use an LLM judge (using the Ollama `qwen3`) to evaluate whether the answer is correct or not. The obtained accuracy is as follows:
| Embedding Model | QA Model | Accuracy | | Embedding Model | QA Model | Accuracy | Reranker |
|------------------------------|-----------------------------------|-----------| |------------------------------|-----------------------------------|-----------|----------------------|
| Ollama / `mxbai-embed-large` | Ollama / `qwen3` | 0.64 | | Ollama / `mxbai-embed-large` | Ollama / `qwen3` | 0.64 | None |
| Ollama / `mxbai-embed-large` | Anthropic / `Claude Sonnet 3.7` | 0.79 | | Ollama / `mxbai-embed-large` | Ollama / `qwen3` | 0.72 | mxbai-rerank-base-v2 |
| Ollama / `mxbai-embed-large` | Anthropic / `Claude Sonnet 3.7` | 0.79 | None |

View file

@ -6,7 +6,7 @@ Your process:
2. Search with specific keywords and phrases from the user's question 2. Search with specific keywords and phrases from the user's question
3. Review the search results and their relevance scores 3. Review the search results and their relevance scores
4. If you need additional context, perform follow-up searches with different keywords 4. If you need additional context, perform follow-up searches with different keywords
5. Provide a comprehensive answer based only on the retrieved documents 5. Provide a short and to the point comprehensive answer based only on the retrieved documents
Guidelines: Guidelines:
- Base your answers strictly on the provided document content - Base your answers strictly on the provided document content
@ -15,6 +15,7 @@ Guidelines:
- Indicate when information is incomplete or when you need to search for additional context - Indicate when information is incomplete or when you need to search for additional context
- If the retrieved documents don't contain sufficient information, clearly state: "I cannot find enough information in the knowledge base to answer this question." - If the retrieved documents don't contain sufficient information, clearly state: "I cannot find enough information in the knowledge base to answer this question."
- For complex questions, consider breaking them down and performing multiple searches - For complex questions, consider breaking them down and performing multiple searches
- Stick to the answer, do not ellaborate or provde context unless asked for it.
Be concise, and always maintain accuracy over completeness. Prefer short, direct answers that are well-supported by the documents. Be concise, and always maintain accuracy over completeness. Prefer short, direct answers that are well-supported by the documents.
""" """

View file

@ -35,23 +35,35 @@ class LLMJudge:
- score: str rating from 1-5 - score: str rating from 1-5
""" """
prompt = f""" prompt = f"""You are an expert evaluator determining whether two answers to the same question are semantically equivalent.
You are an expert judge evaluating the equivalence of two answers to the same question.
Question: {question} QUESTION: {question}
Generated Answer: {answer} GENERATED ANSWER: {answer}
Expected Answer: {expected_answer} EXPECTED ANSWER: {expected_answer}
Your task is to determine if these two answers are equivalent in meaning and both correctly answer the question. Consider: EVALUATION CRITERIA:
Rate as EQUIVALENT (true) if:
Both answers contain the same core factual information
Both directly address the question asked
The key claims and conclusions are consistent
Any additional detail in one answer doesn't contradict the other
1. Do both answers provide the same answer? Rate as NOT EQUIVALENT (false) if:
2. Do both answers directly address the question asked? Factual contradictions exist between the answers
3. Minor differences in wording or style are acceptable if the meaning of the answer is the same. One answer fails to address the core question
4. If one answer is more detailed but the other is correct, they can still be considered equivalent. Key information is missing from one answer that changes the meaning
The answers lead to different conclusions or implications
Be strict but fair in your evaluation. Focus on factual correctness and whether both answers would satisfy someone asking the question.""" GUIDELINES:
- Ignore minor differences in phrasing, style, or formatting
- Focus on semantic meaning rather than exact wording
- Consider both answers correct if they convey the same essential information
- Be tolerant of different levels of detail if the core answer is preserved
- Evaluate based on what a person asking this question would need to know
Respond with JSON containing only: {{"equivalent": true}} or {{"equivalent": false}}"""
response = await self.client.chat( response = await self.client.chat(
model=self.model, model=self.model,