haiku.rag/evaluations
2025-09-30 11:31:51 +03:00
..
datasets Refactor evaluations so that we can perform with multiple datasets.. 2025-09-30 11:31:50 +03:00
__init__.py Refactor evaluations so that we can perform with multiple datasets.. 2025-09-30 11:31:50 +03:00
benchmark.py Add option to skip db in evals 2025-09-30 11:31:51 +03:00
config.py Refactor evaluations so that we can perform with multiple datasets.. 2025-09-30 11:31:50 +03:00
llm_judge.py Use gpt-oss for evaluation LLMJudge, allow it to retry if it fails 2025-09-30 11:31:51 +03:00