Commit graph

62 commits

Author SHA1 Message Date
Yiorgis Gozadinos
e881ce285d
update orb benchmark 2026-05-21 14:23:41 +03:00
Yiorgis Gozadinos
0ee00334c7
drop prose emdashes and semicolons; minor fixes 2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
f1f2e24639
docs: benchmarks — skill-only framing, move inactive datasets to bottom 2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
910d178b00
docs: rebuild Skills section, drop pitch prose 2026-05-20 14:27:16 +03:00
Yiorgis Gozadinos
6f95e2bc27
Delete the standalone QA and analysis agents 2026-05-19 11:39:20 +03:00
Yiorgis Gozadinos
f588e1150b
add vllm Gemma-4-26B skill QA row to wix table 2026-05-18 09:43:13 +03:00
Yiorgis Gozadinos
abc5a210a6
update benchmarks 2026-05-15 10:04:17 +03:00
Yiorgis Gozadinos
d76c645862
Update benchmarks for orb for gemma4 2026-05-14 14:03:19 +03:00
Yiorgis Gozadinos
2189002974
Benchmark update 2026-05-12 17:20:48 +03:00
Yiorgis Gozadinos
cc59dc7203
Use upload_large_folder for hf upload 2026-05-06 13:12:08 +03:00
Yiorgis Gozadinos
ce7201271f
split open_rag_bench dataset into orb_text and orb_multimodal variants 2026-05-06 12:49:03 +03:00
Yiorgis Gozadinos
a45d82f206
Update benchmark 2026-05-05 13:58:48 +03:00
Yiorgis Gozadinos
5c3a864a0b
rename open_rag_bench eval DBs to text/multimodal variants 2026-05-05 12:55:06 +03:00
Yiorgis Gozadinos
17e36fc069
Add first openrag results 2026-05-05 09:50:22 +03:00
Yiorgis Gozadinos
d509b0662d
update benchmarks 2026-04-30 16:54:17 +03:00
Yiorgis Gozadinos
3391335e3a
Update qa wix benchmarks 2026-04-29 14:21:40 +03:00
Yiorgis Gozadinos
318266e681
Update benchmark 2026-04-29 13:27:32 +03:00
Yiorgis Gozadinos
8962987ff9
pin judge model to ollama:qwen3.6 2026-04-29 12:24:06 +03:00
Yiorgis Gozadinos
5e31c15907
docs 2026-04-28 14:44:27 +03:00
Yiorgis Gozadinos
fd327996b8
Configurable judge and reflect models for evaluations 2026-03-18 17:14:11 +02:00
Yiorgis Gozadinos
37f28ea8de
Clean up stale references and dead code 2026-02-20 16:37:15 +02:00
Yiorgis Gozadinos
a889883cfc
Update wix benchmarks 2026-02-10 12:49:24 +02:00
Yiorgis Gozadinos
5f6488e116
Update benchmarks 2026-01-30 13:05:47 +02:00
Yiorgis Gozadinos
2a99cb09bf
Evaluate using ask --deep 2026-01-29 14:22:18 +02:00
Yiorgis Gozadinos
369b3a4bf7
No zip 2026-01-26 16:28:19 +02:00
Yiorgis Gozadinos
9b166969e2
Introduce huggingface dataset for sharing evaluation dbs 2026-01-26 15:53:37 +02:00
Yiorgis Gozadinos
50a1a6171c
Update orb qa accuracy 2026-01-26 10:54:36 +02:00
Yiorgis Gozadinos
756285e91c
Add retrieval benchmarks for ORB 2026-01-22 15:01:37 +02:00
Yiorgis Gozadinos
9bf3a83b5d
Fix typo 2026-01-05 16:39:36 +02:00
Yiorgis Gozadinos
d77bf89aad
Update retrieval/qa benchmark for hotpotqa 2025-12-17 08:27:40 +02:00
Yiorgis Gozadinos
f726229f60
Reformat benchmarks and add hotpotqa placeholder 2025-12-12 12:33:31 +02:00
Yiorgis Gozadinos
be0e3c3472
Update benchmarks 2025-12-10 12:04:07 +02:00
Yiorgis Gozadinos
a1162f8020
rebase from main 2025-12-08 16:07:50 +02:00
Yiorgis Gozadinos
0dc4458070
Fix qa accuracy table for wix 2025-12-06 12:01:13 +02:00
Yiorgis Gozadinos
5dfa07cb2c
Update benchmarks 2025-11-28 12:10:08 +02:00
Yiorgis Gozadinos
4e1752a740
Default eval db location, evaluation script 2025-11-25 11:07:20 +02:00
Yiorgis Gozadinos
b5892699a0
Use Mean Reciprocal Rank for single document evaluation as metric. Use Mean Average Precision for variable document evaluation as metric 2025-11-24 14:49:30 +02:00
Yiorgis Gozadinos
fe14e73514
Benchmarks for win qwen3-embedding:4b mxbai-rerank-base-v2 2025-11-21 18:57:07 +02:00
Yiorgis Gozadinos
9a1e7f211c
Update benchmarks 2025-11-21 15:31:36 +02:00
Yiorgis Gozadinos
63b374c7e9
Name evaluation runs. Use only LLMJudge evaluator. 2025-11-21 13:26:39 +02:00
Yiorgis Gozadinos
dcba9d9a0e
Benchmarks for qwen3-embedding:0.6b 2025-11-14 15:36:36 +02:00
Yiorgis Gozadinos
2f9c907031
Restructure into uv workspace to support minimal and full installations 2025-11-04 17:59:12 +02:00
Yiorgis Gozadinos
a7fd41d24f
Document passing config files when running evaluations 2025-10-30 15:17:02 +02:00
Yiorgis Gozadinos
6f181e07ff
Add Success@K metrics to retrieval benchmarks 2025-10-06 14:54:17 +03:00
Yiorgis Gozadinos
43b9cd50ba
Move evaluations to src/ 2025-09-30 13:21:53 +03:00
Yiorgis Gozadinos
06dbdbd474
Update docs for wix benchmarks 2025-09-30 11:31:52 +03:00
Yiorgis Gozadinos
6f20d9fd6d
Update recall of qwen3 embeddings with mxbai reranking 2025-09-30 11:31:51 +03:00
Yiorgis Gozadinos
842e166041
Adapt how we measure recall when using datasets with multiple sources 2025-09-30 11:31:51 +03:00
Yiorgis Gozadinos
57e86d00fd
Refactor evaluations so that we can perform with multiple datasets..
Introduce Wix dataset.
2025-09-30 11:31:50 +03:00
Yiorgis Gozadinos
87f28ed96b
Update docs 2025-09-30 11:30:35 +03:00